On This Page

Home / Search/ Troubleshoot/Troubleshoot Lakehouse Engines

Troubleshoot Lakehouse Engines ​

Diagnose and resolve common issues with data you ingest into lakehouse engines.


Highlights ​
  • Use Live Data to check each ingest stage: arrival at the Source, Datatyping, and Dataset routing.
  • Fix a Datatype in order: data format first, then field names, then timestamp extraction.
  • Anchor timestamp extraction right before the event time, so Auto doesn’t pick up other numbers.

Lakehouse Engine Troubleshooting Path ​

If your ingested logs don’t look the way you expect, check each stage of ingestion in Live Data, in this order:

  1. Do events arrive? Capture At Source. If you see no events, check your Source configuration, and confirm that the lakehouse engine’s status is Ready.
  2. Do events get the right Datatype? Capture After Datatyping and check the datatype field. If events are Uncategorized or get the wrong Datatype, see Events Get No Datatype or the Wrong One.
  3. Do events have the right fields? If not, see Fields Have Generic Names or Wrong Values.
  4. Does _time match the time the event happened? If not, see Timestamp Issues.
  5. Do events land in the right Dataset? Capture After Dataset detection. If events are Orphaned, see Find Events with No Dataset (Orphaned).

Change one Datatype at a time. Before you move to the next one, verify your change in Live Data with a small sample of events.

Datatyping Issues ​

This section covers issues with how Cribl Search assigns Datatypes and parses events.

To change a stock v2 Datatype, clone it first.

Events Get No Datatype or the Wrong One ​

Events show Uncategorized in the datatype field, or a Datatype other than the one your Datatype rule points to.

Possible Cause ​

Datatype rules run top-down, and the first match wins. Events that match no rule fall through to Auto-Datatyping, and events that Auto-Datatyping doesn’t recognize become Uncategorized.

Common causes:

  • The rule’s expression doesn’t match the events, for example because of a typo in a value it compares against.
  • A broader rule above your rule matches the events first.
  • The rule isn’t enabled.
Recommendations ​
  1. To see which Datatypes your stored events got, run a query of this form:

    dataset="<your_dataset>" | summarize count() by datatype
  2. Compare the rule’s expression with the events you capture At Source in Live Data. Point the expression at fields that exist before Datatyping, such as _raw and __inputId. See Datatype Rule Expressions.

  3. Drag more specific rules above broader ones, and make sure that Enabled is checked on each rule.

Fields Have Generic Names or Wrong Values ​

Events have fields named field1, field2, and so on, or field values hold fragments of the original event.

Possible Cause ​
  • The Datatype’s Data format doesn’t match the content of _raw. For example, a Delimited Text Datatype splits a JSON event at every comma.
  • A Delimited Text Datatype has an empty Optional field list, so Cribl Search names the fields field1, field2, and so on.
Recommendations ​
  1. Capture the events At Source in Live Data, and check the format of _raw.
  2. In the Datatype, set Data format to match it. For example, use JSON Newline Delimited for one JSON object per line. See Set the Data Format.
  3. For Delimited Text, list the field names in Optional field list, in the order they appear in _raw.

Fix the data format before you adjust timestamp extraction. When Time field names a parsed field, timestamp extraction depends on correct parsing.

Timestamp Issues ​

This section covers events whose _time doesn’t match the time the event happened.

When a Datatype’s Extraction type is Auto, Cribl Search finds the first match of Timestamp anchor in the Time field. Starting right after that match, it scans up to Scan depth characters for anything that looks like a timestamp. With the default anchor (/^/) and scan depth (150), the scan starts at the beginning of the field. If that range holds other dates or numbers, such as expiration dates, message IDs, or cookie values, Auto can pick one of them instead of the event time.

Auto reads a standalone number as Unix time when it has 10 digits and starts with 1, optionally followed by three more digits for milliseconds. Such numbers fall between September 2001 and May 2033, so an ID or counter in this format looks like a valid timestamp.

Cribl Search then checks the extracted timestamp against two sets of bounds:

  • The Datatype’s Earliest timestamp allowed and Future timestamp allowed, which default to 0 (January 1, 1970, in Unix time) and +10years. Cribl Search sets timestamps outside this range to the nearest bound. These bounds don’t apply when Time field holds a number, which Cribl Search uses directly as Unix time, in seconds or milliseconds. See Set Timestamp Extraction.
  • The Dataset’s Expected Time Range. For logs outside this range, Cribl Search resets _time to now() and keeps the original value in _original_time.

Events Have Timestamps in the Future ​

Events carry an _original_time field whose value lies in the future. With the default Future timestamp allowed, that value is up to 10 years ahead.

Possible Cause ​

Auto picked up a token that isn’t the event time. For example, an expiration or schedule date, or an ID that it read as Unix time. Because the result fell after the Dataset’s Latest expected timestamp, Cribl Search reset _time to the ingest time and kept the extracted value in _original_time.

Recommendations ​
  1. To find the affected Datatypes, run a query of this form:

    dataset="<your_dataset>" | where isnotnull(_original_time) | summarize count() by datatype
  2. Capture the affected events At Source in Live Data, and find where the real event time appears in _raw.

  3. Narrow down where Auto looks. See Point Timestamp Extraction at the Event Time.

Events Have Timestamps in the Past ​

Live Data shows events arriving, but your searches over recent time ranges don’t return them. When you capture the events After Datatyping, their _time lies years in the past, sometimes near January 1, 1970.

Possible Cause ​
  • Auto read an ID or counter in a cookie, header, or other field as Unix time, which lands it in the past.
  • Time field names a field with a numeric value. Cribl Search uses that number as Unix time, so a small value lands near 1970. This usually points to an issue in the source data rather than in your Datatype.

By default, no bound catches these timestamps. Earliest timestamp allowed defaults to 0, which allows any time since 1970, and new Datasets have no Earliest expected timestamp.

Recommendations ​
  1. Narrow down where Auto looks. See Point Timestamp Extraction at the Event Time.
  2. If the source sends a wrong value, fix it at the source.
  3. To keep outliers searchable, set Earliest expected timestamp on the Dataset. Cribl Search then resets _time of older logs to now() and keeps the original value in _original_time. See Expected Time Range.

Datatype Timestamp Settings Have No Effect ​

You change the timestamp extraction settings of a Datatype, but _time on new events doesn’t change.

Possible Cause ​
  • Events arrive with a numeric _time already set, for example by Cribl Stream. For events that get a Datatype, Cribl Search keeps that value and skips timestamp extraction. See Override Cribl Search Processing.
  • Events don’t get this Datatype. For Uncategorized events, Cribl Search runs Auto over the first 150 characters of _raw instead of your Datatype settings. See Events Get No Datatype or the Wrong One.
  • Time field names a field that doesn’t exist when Cribl Search extracts the timestamp. In that case, Cribl Search scans _raw instead, without a warning. This happens when the field:
    • Has a typo in its name, or is a nested path such as event.time.
    • Comes from Add fields to events or Schema Maps, which Cribl Search applies after timestamp extraction.
    • Is _time, and events arrive without it.
  • Timestamp anchor doesn’t match the event. In that case, Cribl Search scans from the beginning of the field.
Recommendations ​
  • If _time comes from upstream, fix it there, or remove it upstream so the Datatype extracts the timestamp.
  • Set Time field to _raw, or to a top-level field that Data format or Additional Extractions produce. Cribl Search parses these fields before it extracts the timestamp. Don’t set Time field to _time.
  • Test your Timestamp anchor regex against a sample _raw from Live Data.

Point Timestamp Extraction at the Event Time ​

To make Auto find the right timestamp, narrow down where it looks. In the Datatype’s timestamp extraction settings:

  1. Choose where to look:

    • If Data format parses the event time into its own top-level field, such as a JSON key, set Time field to that field.
    • Otherwise, set Time field to _raw.
  2. Set Timestamp anchor to a regex that matches the text right before the timestamp. For example:

    Where the timestamp appearsTimestamp anchor
    At the start of the line/^/ (default)
    In a JSON key, as in "CreationTime":"2026-..."/"CreationTime":\s*"/
    In brackets, as in Apache access logs/\[/
    In the third column of a comma-separated line/^(?:[^,]*,){2}/
  3. Lower Scan depth so the scan covers the timestamp and stops before free text, URLs, or IDs. For example, an ISO 8601 timestamp with milliseconds and a time zone offset fits in 30 characters.

  4. If Auto still picks the wrong value, set Extraction type to Manual, and enter the exact strptime format of your timestamps.

  5. Narrow Earliest timestamp allowed and Future timestamp allowed to a plausible range for your data. Many events with a timestamp exactly at one of these bounds point to an extraction issue. These bounds don’t apply to numeric Time field values.

  6. Save the Datatype. Capture a few new events After Datatyping in Live Data, and check that _time matches the event time.