Datasets in ES|QL Data Federation
A dataset makes a named collection of files in external storage available to ES|QL. It records which connected data source and files to use, together with any settings or schema definitions needed to interpret them. You can query the dataset by name without ingesting its data into Elasticsearch.
For the overall mental model, from connecting external storage to querying a dataset, refer to how ES|QL Data Federation works.
This feature is experimental. It is not intended for production use and there are no guarantees around performance, scale, or stability in this release.
To define a dataset, work through the following decisions:
- Select a data source. Use a connected data source that provides access to the external storage. One data source can serve multiple datasets.
- Name the dataset. The name identifies the dataset in an ES|QL query. Dataset names share a namespace with indices, data streams, aliases, and views, so a dataset cannot use the name of any of these existing objects.
- Select the files. Use a storage URI and resource pattern to select files in one supported file format.
- Determine the schema. Let Elasticsearch infer and reconcile the schema, or declare column names and data types explicitly.
- Adjust dataset behavior. Add dataset settings when you need to change the defaults for file discovery, parsing, error handling, schema resolution, or query parallelism.
- Describe the dataset. Add an optional description to explain what the dataset contains or how it is used.
After defining what the dataset reads and how to interpret it, create and manage the dataset in Kibana or with the /_query/dataset API.
Reference the dataset by name in the ES|QL FROM command. To learn how queries discover files and reduce the amount of external data read, refer to Query data with ES|QL Data Federation.