Load historical data into a time series data stream

There are two methods for loading historical documents into a time series data stream (TSDS). If the timestamps fall inside the eligible write window, turn on past index creation and load the documents into the existing stream. If they're older, load them into a separate historical stream.

Follow Load data within the eligible write window or Load data beyond the eligible write window based on whether your timestamps fall inside that window.

Loading months of historical data can trigger significant storage use, force merge activity, and lifecycle processing in parallel. Verify that your cluster has enough available resources before you start.

This approach works well when you're backfilling recent history alongside live ingestion, such as late-arriving metrics or a short bootstrap period.

  1. Turn on past index creation

    Elasticsearch can create missing backing indices when you add data that precedes existing time ranges. To enable this feature, update the cluster settings:

    				PUT _cluster/settings
    					{
      "persistent": {
        "data_stream.past_tsdb_index_creation_enabled": true,
        "data_streams.past_tsdb_index_interval": "2d"
      }
    }
    		
    1. By default, each past backing index covers one day of data. Refer to data_streams.past_tsdb_index_interval.

    After you turn on past index creation, Elasticsearch creates past backing indices as documents arrive. Write-time deduplication and TSDS storage optimizations apply to historical data the same way they apply to live data.

    Note

    You need the auto_configure index privilege to trigger past index creation. For details, refer to Secure a TSDS.

  2. Index the historical documents

    Point your migration or replay pipeline at the live time series data stream. You can use the same APIs you use for live data.

    If the stream already has a downsampling lifecycle, those past indices might qualify immediately. Elasticsearch ages them from the data they contain, not from when the index was created. To limit concurrent downsampling per data stream, configure the data_streams.lifecycle.downsampling.max_indices_in_progress cluster setting.

    For an example of setting up a TSDS and loading historical data into it, refer to Set up a time series data stream.

  3. Confirm the load

    Use the get data stream API to check that backing indices cover the timestamps you loaded. For example:

    				GET _data_stream/metrics-weather-sensors
    		

    The response lists each backing index and the time range it accepts.

You can't load data older than the eligible write window directly into a TSDS. For example, if downsampling makes indices read-only after seven days, you can't backfill eighteen months of history into that same data stream.

Instead, create a separate historical TSDS without a lifecycle, load the data, then add a data stream lifecycle when the load is complete.

  1. Create an index template for the historical data stream

    Use the same mappings as your live TSDS, but don't include a lifecycle policy in the template. For example, use the create index template API:

    				PUT _index_template/metrics-historical
    					{
      "index_patterns": ["metrics-historical-*"],
      "data_stream": {},
      "template": {
        "settings": {
          "index.mode": "time_series"
        },
        "mappings": {
          "properties": {
            "@timestamp": { "type": "date" },
            "sensor_id": { "type": "keyword", "time_series_dimension": true },
            "temperature": { "type": "half_float", "time_series_metric": "gauge" }
          }
        }
      }
    }
    		
  2. Create the historical data stream

    Create a data stream with a name that matches the pattern in the index template. For example, use the create a data stream API:

    				PUT _data_stream/metrics-historical-2024
    		
  3. Index historical data

    Index historical data into the historical data stream while current data continues flowing into the original TSDS.

    Important

    Historical data must fit on the target tier as a whole before you enable data stream lifecycle. If you're importing a large data set, split it into batches. Each batch should fit within available disk space at indexing time.

    For an example of how to check disk space with the cat allocation API, refer to Estimate the amount of required disk capacity.

  4. Add data stream lifecycle

    When the load is complete, add a data stream lifecycle to the historical data stream. For example, use the update data stream lifecycles API:

    				PUT _data_stream/metrics-historical-2024/_lifecycle
    					{
      "enabled": true,
      "data_retention": "365d",
      "downsampling": [
        {
          "after": "7d",
          "fixed_interval": "10m"
        }
      ]
    }
    		

    Processing begins immediately and creates a backlog of downsampling work. When you add a lifecycle to a data stream with many indices that qualify for downsampling, data stream lifecycle can queue multiple downsampling operations at once. To limit concurrent downsampling per data stream, configure the data_streams.lifecycle.downsampling.max_indices_in_progress cluster setting. For details, refer to Downsample with a data stream lifecycle. If you include data_retention settings, data stream lifecycle deletes expired backing indices but does not remove the data stream itself.

  5. Query across both data streams

    Query both streams with a wildcard pattern or a data stream alias. For example, use the search API:

    				GET metrics-*/_search
    					{
      "size": 10,
      "sort": [{ "@timestamp": "desc" }]
    }
    		

    The results include documents from both the live stream and the historical stream.

Delete historical data streams manually when their data is no longer needed.

Backfill and creation of past indices have the following limitations:

  • System data streams are excluded.
  • Cross-cluster replication (CCR) follower data streams rely on the leader data stream, so you can't backfill follower streams directly.