Marktechpost iconMarktechpostSep 5, 2026

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

A single datasets.invent call sets domains, row count, output format, and language expansion, and the rows download as JSONL, JSON, CSV, or Parquet. The dataset ID then passes straight into AutoScientist, closing an intent-to-trained-model loop.

Adaption Labs Introduces ‘Invent a Dataset’: Training Data Generated From a Task Description, Not a Seed Corpus

Share this story

Send the public story page.

Useful takeaways from this story.

A single datasets.invent call sets domains, row count, output format, and language expansion, and the rows download as JSONL, JSON, CSV, or Parquet.

The dataset ID then passes straight into AutoScientist, closing an intent-to-trained-model loop.

There is no seed corpus, no schema design, and no labeling guide.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

A single datasets.invent call sets domains, row count, output format, and language expansion, and the rows download as JSONL, JSON, CSV, or Parquet. The dataset ID then passes straight into AutoScientist, closing an intent-to-trained-model loop. There is no seed corpus, no schema design, and no labeling guide.

How it works

Details worth keeping

There is no seed corpus, no schema design, and no labeling guide.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app