The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
Error code: DatasetGenerationCastError
Exception: DatasetGenerationCastError
Message: An error occurred while generating the dataset
All the data files must have the same columns, but at some point there are 4 new columns ({'signup_date', 'customer_id', 'acquisition_channel', 'region'}) and 3 missing columns ({'date', 'spend', 'channel'}).
This happened while the csv dataset builder was generating data using
hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark/customers.csv (at revision 4536d2e001f407cf37d4b607070dedeb0bac4ddd), ['hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/ad_spend.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/customers.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/inventory.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/orders.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/products.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/returns.csv']
Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
writer.write_table(table)
~~~~~~~~~~~~~~~~~~^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 765, in write_table
self._write_table(pa_table, writer_batch_size=writer_batch_size)
~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 773, in _write_table
pa_table = table_cast(pa_table, self._schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2378, in table_cast
return cast_table_to_schema(table, schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2306, in cast_table_to_schema
raise CastError(
...<3 lines>...
)
datasets.table.CastError: Couldn't cast
customer_id: string
signup_date: string
region: string
acquisition_channel: string
-- schema metadata --
pandas: '{"index_columns": [{"kind": "range", "name": null, "start": 0, "' + 772
to
{'date': Value('string'), 'channel': Value('string'), 'spend': Value('float64')}
because column names don't match
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
~~~~~~~~~~~~~~~~~~~~~~~~~^
builder, max_dataset_size_bytes=max_dataset_size_bytes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
for job_id, done, content in self._prepare_split_single(
~~~~~~~~~~~~~~~~~~~~~~~~~~^
gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
):
^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1850, in _prepare_split_single
raise DatasetGenerationCastError.from_cast_error(
...<4 lines>...
)
datasets.exceptions.DatasetGenerationCastError: An error occurred while generating the dataset
All the data files must have the same columns, but at some point there are 4 new columns ({'signup_date', 'customer_id', 'acquisition_channel', 'region'}) and 3 missing columns ({'date', 'spend', 'channel'}).
This happened while the csv dataset builder was generating data using
hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark/customers.csv (at revision 4536d2e001f407cf37d4b607070dedeb0bac4ddd), ['hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/ad_spend.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/customers.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/inventory.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/orders.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/products.csv', 'hf://datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark@4536d2e001f407cf37d4b607070dedeb0bac4ddd/returns.csv']
Please either edit the data files to have matching columns, or separate them into different configurations (see docs at https://hf.co/docs/hub/datasets-manual-configuration#multiple-configurations)Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
date string | channel string | spend float64 |
|---|---|---|
2026-01-01 | Paid Search | 615.07 |
2026-01-01 | Social | 314.36 |
2026-01-01 | Display | 267.71 |
2026-01-01 | Affiliate | 186.98 |
2026-01-02 | Paid Search | 670.06 |
2026-01-02 | Social | 381.39 |
2026-01-02 | Display | 213.85 |
2026-01-02 | Affiliate | 191.72 |
2026-01-03 | Paid Search | 510.66 |
2026-01-03 | Social | 339.72 |
2026-01-03 | Display | 260.69 |
2026-01-03 | Affiliate | 164.09 |
2026-01-04 | Paid Search | 457.81 |
2026-01-04 | Social | 419.06 |
2026-01-04 | Display | 292.14 |
2026-01-04 | Affiliate | 188.37 |
2026-01-05 | Paid Search | 506.25 |
2026-01-05 | Social | 442.63 |
2026-01-05 | Display | 223.03 |
2026-01-05 | Affiliate | 181.11 |
2026-01-06 | Paid Search | 669 |
2026-01-06 | Social | 366.65 |
2026-01-06 | Display | 283.37 |
2026-01-06 | Affiliate | 174.36 |
2026-01-07 | Paid Search | 636.52 |
2026-01-07 | Social | 309.51 |
2026-01-07 | Display | 301.38 |
2026-01-07 | Affiliate | 173.83 |
2026-01-08 | Paid Search | 585.91 |
2026-01-08 | Social | 424.02 |
2026-01-08 | Display | 276.55 |
2026-01-08 | Affiliate | 178.35 |
2026-01-09 | Paid Search | 604.04 |
2026-01-09 | Social | 352.96 |
2026-01-09 | Display | 272.41 |
2026-01-09 | Affiliate | 219.84 |
2026-01-10 | Paid Search | 486.26 |
2026-01-10 | Social | 416.54 |
2026-01-10 | Display | 306.76 |
2026-01-10 | Affiliate | 145.69 |
2026-01-11 | Paid Search | 656.17 |
2026-01-11 | Social | 324.54 |
2026-01-11 | Display | 267.95 |
2026-01-11 | Affiliate | 148.83 |
2026-01-12 | Paid Search | 632.72 |
2026-01-12 | Social | 364.79 |
2026-01-12 | Display | 277.27 |
2026-01-12 | Affiliate | 175.62 |
2026-01-13 | Paid Search | 487.63 |
2026-01-13 | Social | 459.75 |
2026-01-13 | Display | 257.6 |
2026-01-13 | Affiliate | 204.31 |
2026-01-14 | Paid Search | 642.68 |
2026-01-14 | Social | 452.32 |
2026-01-14 | Display | 207.07 |
2026-01-14 | Affiliate | 211.11 |
2026-01-15 | Paid Search | 476.16 |
2026-01-15 | Social | 487.45 |
2026-01-15 | Display | 273.55 |
2026-01-15 | Affiliate | 163.49 |
2026-01-16 | Paid Search | 485.37 |
2026-01-16 | Social | 435.54 |
2026-01-16 | Display | 203.38 |
2026-01-16 | Affiliate | 154 |
2026-01-17 | Paid Search | 533.29 |
2026-01-17 | Social | 447.39 |
2026-01-17 | Display | 242.56 |
2026-01-17 | Affiliate | 166.41 |
2026-01-18 | Paid Search | 446.58 |
2026-01-18 | Social | 438.99 |
2026-01-18 | Display | 214.62 |
2026-01-18 | Affiliate | 169.19 |
2026-01-19 | Paid Search | 583.72 |
2026-01-19 | Social | 451.26 |
2026-01-19 | Display | 198.23 |
2026-01-19 | Affiliate | 137.9 |
2026-01-20 | Paid Search | 562.08 |
2026-01-20 | Social | 445.86 |
2026-01-20 | Display | 297.13 |
2026-01-20 | Affiliate | 174.66 |
2026-01-21 | Paid Search | 657.39 |
2026-01-21 | Social | 441.96 |
2026-01-21 | Display | 239.62 |
2026-01-21 | Affiliate | 195.06 |
2026-01-22 | Paid Search | 630.02 |
2026-01-22 | Social | 395.97 |
2026-01-22 | Display | 227.07 |
2026-01-22 | Affiliate | 199.11 |
2026-01-23 | Paid Search | 605.15 |
2026-01-23 | Social | 402.68 |
2026-01-23 | Display | 253.24 |
2026-01-23 | Affiliate | 221.36 |
2026-01-24 | Paid Search | 556.22 |
2026-01-24 | Social | 403 |
2026-01-24 | Display | 307.67 |
2026-01-24 | Affiliate | 195.64 |
2026-01-25 | Paid Search | 589.79 |
2026-01-25 | Social | 476.84 |
2026-01-25 | Display | 232.66 |
2026-01-25 | Affiliate | 167.61 |
E-commerce AI Data Analyst Agent Benchmark
A synthetic e-commerce dataset for evaluating AI data analyst agents on realistic, multi-step business analysis, data-quality investigation, and analytical reasoning.
This dataset is part of the E-commerce AI Data Analyst Agent Benchmark.
Dataset summary
This dataset supports evaluation of AI data analyst agents on realistic, multi-step e-commerce analysis.
It contains:
customers.csvproducts.csvorders.csvreturns.csvad_spend.csvinventory.csv
The benchmark covers tasks involving:
- data retrieval and aggregation
- business analysis
- data-quality investigation
- reconciliation
- analytical reasoning
- business interpretation
Data provenance
The datasets are synthetic and were generated specifically for this benchmark. They are not sourced from a real company's customer or transaction database.
The data intentionally contains controlled defects for some benchmark tasks, including duplicate records, missing or orphan references, invalid numeric values, and return-linkage problems.
Intended use
This dataset is intended for:
- AI-agent evaluation
- data analyst agent benchmarking
- research and experimentation
- reproducible testing of analytical workflows
- educational and noncommercial projects
Limitations
This is a synthetic benchmark dataset. Results obtained using it should not be interpreted as evidence of performance on any particular real-world company or production dataset.
The dataset is designed for benchmark evaluation and contains deliberately constructed data-quality defects. Those defects should not be treated as representative estimates of real-world data quality.
Related benchmark
The complete benchmark, including tasks, ground truth, evaluator, adversarial cases, calibration tests, pilot results, and documentation is available in the GitHub repository:
https://github.com/Omcrec/ecommerce-ai-data-analyst-agent-benchmark
License
The dataset and associated benchmark content in this repository are licensed under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0).
You may copy, modify, remix, and redistribute the material for noncommercial purposes with appropriate attribution.
Commercial use, commercial distribution, or sale of the dataset or substantial portions of its content is not permitted under this license.
License: https://creativecommons.org/licenses/by-nc/4.0/
Version
Benchmark version: 1.0.0
- Downloads last month
- 98