﻿---
title: Load historical data into a time series data stream
description: Load historical documents into a time series data stream using the live stream for recent timestamps and a separate stream for data older than the eligible write window.
url: https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/load-historical-tsds
products:
  - Elastic Documentation
  - Elasticsearch
applies_to:
  - Elastic Stack: Generally available since 9.5
---

# Load historical data into a time series data stream
There are two methods for loading historical documents into a time series data stream (TSDS).
If the timestamps fall inside the [eligible write window](/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/time-bound-tsds#tsds-past-index-creation), turn on past index creation and load the documents into the existing stream.
If they're older, load them into a separate historical stream.
Follow [Load data within the eligible write window](#load-data-within-the-eligible-write-window) or [Load data beyond the eligible write window](#load-data-beyond-the-eligible-write-window) based on whether your timestamps fall inside that window.

## Before you begin

Loading months of historical data can trigger significant storage use, force merge activity, and lifecycle processing in parallel.
Verify that your cluster has enough available resources before you start.

## Load data within the eligible write window

This approach works well when you're backfilling recent history alongside live ingestion, such as late-arriving metrics or a short bootstrap period.
<stepper>
  <step title="Turn on past index creation">
    Elasticsearch can create missing backing indices when you add data that precedes existing time ranges.
    To enable this feature, update the cluster settings:
    ```json

    {
      "persistent": {
        "data_stream.past_tsdb_index_creation_enabled": true,
        "data_streams.past_tsdb_index_interval": "2d" <1>
      }
    }
    ```

    1. By default, each past backing index covers one day of data. Refer to [`data_streams.past_tsdb_index_interval`](https://docs-v3-preview.elastic.dev/elastic/docs-builder/docs/4384/reference/elasticsearch/configuration-reference/miscellaneous-cluster-settings#time-series-data-stream).
    After you turn on past index creation, Elasticsearch creates past backing indices as documents arrive.
    Write-time deduplication and TSDS storage optimizations apply to historical data the same way they apply to live data.
    <note>
      You need the `auto_configure` index privilege to trigger past index creation.
      For details, refer to [Secure a TSDS](/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/set-up-tsds#secure-tsds).
    </note>
  </step>

  <step title="Index the historical documents">
    Point your migration or replay pipeline at the live time series data stream.
    You can use the same APIs you use for live data.If the stream already has a downsampling lifecycle, those past indices might qualify immediately.
    Elasticsearch ages them from the data they contain, not from when the index was created.
    To limit concurrent downsampling per data stream, configure the [`data_streams.lifecycle.downsampling.max_indices_in_progress`](https://docs-v3-preview.elastic.dev/elastic/docs-builder/docs/4384/reference/elasticsearch/configuration-reference/data-stream-lifecycle-settings#data-streams-lifecycle-downsampling-max-indices-in-progress) cluster setting.For an example of setting up a TSDS and loading historical data into it, refer to [Set up a time series data stream](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/set-up-tsds).
  </step>

  <step title="Confirm the load">
    Use the [get data stream API](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-get-data-stream) to check that backing indices cover the timestamps you loaded.
    For example:
    ```json
    ```
    The response lists each backing index and the time range it accepts.
  </step>
</stepper>


## Load data beyond the eligible write window

You can't load data older than the eligible write window directly into a TSDS.
For example, if downsampling makes indices read-only after seven days, you can't backfill eighteen months of history into that same data stream.
Instead, create a separate historical TSDS without a lifecycle, load the data, then add a [data stream lifecycle](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/lifecycle/data-stream) when the load is complete.
<stepper>
  <step title="Create an index template for the historical data stream">
    Use the same mappings as your live TSDS, but don't include a lifecycle policy in the template.
    For example, use the [create index template](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-put-index-template) API:
    ```json

    {
      "index_patterns": ["metrics-historical-*"],
      "data_stream": {},
      "template": {
        "settings": {
          "index.mode": "time_series"
        },
        "mappings": {
          "properties": {
            "@timestamp": { "type": "date" },
            "sensor_id": { "type": "keyword", "time_series_dimension": true },
            "temperature": { "type": "half_float", "time_series_metric": "gauge" }
          }
        }
      }
    }
    ```
  </step>

  <step title="Create the historical data stream">
    Create a data stream with a name that matches the pattern in the index template.
    For example, use the [create a data stream](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-create-data-stream) API:
    ```json
    ```
  </step>

  <step title="Index historical data">
    Index historical data into the historical data stream while current data continues flowing into the original TSDS.
    <important>
      Historical data must fit on the target tier as a whole before you enable data stream lifecycle.
      If you're importing a large data set, split it into batches.
      Each batch should fit within available disk space at indexing time.For an example of how to check disk space with the [cat allocation API](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-cat-allocation), refer to [Estimate the amount of required disk capacity](/elastic/docs-builder/docs/4384/troubleshoot/elasticsearch/increase-capacity-data-node#estimate-required-capacity).
    </important>
  </step>

  <step title="Add data stream lifecycle">
    When the load is complete, add a [data stream lifecycle](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/lifecycle/data-stream) to the historical data stream.
    For example, use the [update data stream lifecycles](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-indices-put-data-lifecycle) API:
    ```json

    {
      "enabled": true,
      "data_retention": "365d",
      "downsampling": [
        {
          "after": "7d",
          "fixed_interval": "10m"
        }
      ]
    }
    ```
    Processing begins immediately and creates a backlog of downsampling work.
    When you add a lifecycle to a data stream with many indices that qualify for downsampling, data stream lifecycle can queue multiple downsampling operations at once.
    To limit concurrent downsampling per data stream, configure the [`data_streams.lifecycle.downsampling.max_indices_in_progress`](https://docs-v3-preview.elastic.dev/elastic/docs-builder/docs/4384/reference/elasticsearch/configuration-reference/data-stream-lifecycle-settings#data-streams-lifecycle-downsampling-max-indices-in-progress) cluster setting.
    For details, refer to [Downsample with a data stream lifecycle](/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/run-downsampling#downsample-with-a-data-stream-lifecycle).
    If you include `data_retention` settings, data stream lifecycle deletes expired backing indices but does not remove the data stream itself.
  </step>

  <step title="Query across both data streams">
    Query both streams with a wildcard pattern or a [data stream alias](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/data-store/aliases).
    For example, use the [search](https://www.elastic.co/docs/api/doc/elasticsearch/operation/operation-search) API:
    ```json

    {
      "size": 10,
      "sort": [{ "@timestamp": "desc" }]
    }
    ```
    The results include documents from both the live stream and the historical stream.
  </step>
</stepper>

Delete historical data streams manually when their data is no longer needed.

## Limitations

Backfill and creation of past indices have the following limitations:
- System data streams are excluded.
- Cross-cluster replication (CCR) follower data streams rely on the leader data stream, so you can't backfill follower streams directly.


## Next steps

- [Downsample a time series data stream](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/downsampling-time-series-data-stream) to reduce storage after historical data ages
- [Reindex a time series data stream](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/reindex-tsds) if you need to copy data to a new TSDS instead of backfilling in place


## Related pages

- [Time-bound indices](https://www.elastic.co/elastic/docs-builder/docs/4384/manage-data/data-store/data-streams/time-bound-tsds) for eligible write window and past index creation details