﻿---
title: File formats for ES|QL Data Federation
description: Compare the file formats, extensions, compression codecs, and schema sources supported by ES|QL Data Federation datasets.
url: https://www.elastic.co/elastic/docs-builder/docs/4384/reference/query-languages/esql/esql-data-federation-file-formats
products:
  - Elasticsearch
applies_to:
  - Elastic Cloud Serverless: Unavailable
  - Elastic Stack: Experimental since 9.5
---

# File formats for ES|QL Data Federation
ES|QL Data Federation reads Parquet, newline-delimited JSON (NDJSON), comma-separated values (CSV), and tab-separated values (TSV) files from external storage. Use the file extension, schema source, and compression support in this reference when defining a [dataset](https://www.elastic.co/elastic/docs-builder/docs/4384/reference/query-languages/esql/esql-data-federation-datasets).
<warning>
  This feature is experimental. It is not intended for production use and there are no guarantees around performance, scale, or stability in this release.
</warning>

The following table compares the supported file formats:

| Format  | Recognized extensions                                             | Schema source | Compression                 |
|---------|-------------------------------------------------------------------|---------------|-----------------------------|
| Parquet | `.parquet`<applies-to>Elastic Stack: Planned</applies-to> `.parq` | File metadata | Internal per column chunk   |
| NDJSON  | `.ndjson`, `.jsonl`, `.json`                                      | Sampled rows  | Uncompressed, gzip, or zstd |
| CSV     | `.csv`                                                            | Sampled rows  | Uncompressed, gzip, or zstd |
| TSV     | `.tsv`                                                            | Sampled rows  | Uncompressed, gzip, or zstd |


## Select one format per dataset

Scope each dataset to one file format. Elasticsearch infers the format when the resource pattern implies exactly one registered format. For example, `**/*.parquet`, `_schema.parquet,events/**/*.parquet`, and `a.csv,b.csv.gz` each imply one format. Compression suffixes don't count as a second format.
Set [`format`](/elastic/docs-builder/docs/4384/reference/query-languages/esql/esql-data-federation-dataset-settings#format) explicitly for extensionless resources such as `hits/*` and for mixed patterns such as `*.{parquet,csv}`. Alternatively, use a [resource pattern](https://www.elastic.co/elastic/docs-builder/docs/4384/reference/query-languages/esql/esql-data-federation-patterns) that selects one format, and create another dataset for files in a different format.
<important>
  An explicit `format` selects the reader for every file that the resource pattern matches, including files with unrecognized extensions such as `.log.gz`. A file whose extension maps to a different registered format is rejected rather than skipped. For details, refer to [`format`](/elastic/docs-builder/docs/4384/reference/query-languages/esql/esql-data-federation-dataset-settings#format).
</important>


## Compression

A text file can be uncompressed or compressed with a codec identified by its final extension. For example, `clicks.csv`, `clicks.csv.gz`, and `clicks.csv.zst` are all CSV files.

| Codec        | Extensions      |
|--------------|-----------------|
| Uncompressed | None            |
| gzip         | `.gz`, `.gzip`  |
| zstd         | `.zst`, `.zstd` |

Parquet declares compression internally for each column chunk and does not support whole-file compression. The supported internal codecs are `UNCOMPRESSED`, `SNAPPY`, `ZSTD`, `GZIP`, `LZ4_RAW`, and legacy Hadoop-framed `LZ4` for reading.

## Schema handling

Parquet files include their schema in file metadata. Elasticsearch infers schemas for CSV, TSV, and NDJSON by sampling rows. Refer to [schema inference](https://www.elastic.co/elastic/docs-builder/docs/4384/reference/query-languages/esql/esql-data-federation-schema) to learn how Elasticsearch reconciles schemas across multiple files.