Adaption Labs Turns a Single Prompt Into 20,000 AI Training Samples
Adaption's new Invent a Dataset system generates post-training data from a plain text description, beating five frontier model APIs on quality and diversity.
- Adaption released the technical report for Invent a Dataset, a prompt-to-post-training-data system.
- Beats five frontier APIs with 17% higher quality and 19% greater sample diversity across 8 task types.
- Diversity lead widens to 37% at 20K samples with 0.0% duplicates reported.
- Supports instruction pairs for SFT and preference pairs for DPO, exported as JSONL, JSON, CSV, or Parquet.
- Language expansion offers translate and locale-specific localize modes with sample rate control.
- Dataset IDs feed directly into AutoScientist to close an intent-to-trained-model loop.
Adaption Labs turns behavior prompts into training datasets
Adaption Labs has published a technical report for Invent a Dataset, a hosted API that generates post-training data from a description of the behavior a model should learn. The service targets zero-data projects, where a team has no seed corpus, labeled examples, or production traffic for a new capability.
Post-training covers techniques such as supervised fine-tuning and preference optimization, which adapt a pretrained model to specific tasks or response patterns. These techniques require varied, accurate examples. Creating those examples often involves manual schema design, collection, labeling, filtering, and repeated quality checks.
Scale widens the diversity gap
Adaption compared Invent a Dataset with frontier-model APIs from Anthropic, Google, OpenAI, DeepSeek, and Z.ai. The benchmark covered eight task types and datasets containing as many as 20,000 samples.
| Metric | Reported result |
|---|---|
| Quality | 17% relative gain over competing generators |
| Diversity | 19% relative gain across the benchmark |
| Diversity at 200 samples | Approximately equal to the comparison APIs |
| Diversity at 20,000 samples | 37% relative gain |
| Duplicates at 20,000 samples | 0.0% reported duplicate rate |
Relative gains describe proportional improvement over a baseline, rather than percentage-point increases. The growing diversity advantage suggests that competing generators repeated patterns more often as dataset size increased.
Adaption also fine-tuned models on data from each generator. According to the report, models trained on Invent datasets ranked consistently higher across the tested architectures, connecting the dataset-level scores with downstream training performance.
These results come from an evaluation conducted and published by Adaption. Teams comparing generators should review the report’s task definitions, judging method, baselines, and variance, then reproduce the evaluation with prompts and models that match their intended deployment.
From one prompt to a portable file
The API runs dataset generation asynchronously through a short workflow:
- Call
datasets.inventwith the requested behavior, domain, size, and output type. - Store the returned dataset ID while the status is
running. - Poll
datasets.getuntil the status becomessucceededorfailed. - Download the completed rows as JSONL, JSON, CSV, or Parquet.
The downloaded file can be used outside Adaption’s platform. Two dataset structures cover common post-training methods:
| Output type | Contents | Typical use |
|---|---|---|
instruction_dataset |
Prompt and completion pairs | Supervised fine-tuning |
preference_pairs |
Chosen and rejected completions | Preference training such as direct preference optimization, or DPO |
Domains, cost estimates, and safe retries
Domain codes provide the main control over subject matter. Developers can retrieve available values with datasets.invent_domains, then select a broad domain such as medical or a narrower code such as medical.symptoms_diagnosis.
A prompt field of up to 10,000 characters further specifies the desired content and behavior. Setting estimate=True checks the request against the account’s credit balance without starting a billable generation run. An idempotency_key makes retries safer by returning the original dataset for repeated requests with the same key.
Translation with locale controls
Language expansion supports two modes with different levels of adaptation:
translatecreates one variant for each target language.localizecreates variants for country-language pairs and uses locale-specific wording.
A sample_rate from 0.01 to 1 controls the fraction of invented rows expanded into additional languages or locales. Adaption bills for the expanded output rows, so developers should include those variants when estimating cost.
Dataset generation feeds an optimization loop
Invent dataset IDs can pass directly into AutoScientist, Adaption’s system for optimizing both the training data and the training recipe against a specified evaluation objective. This creates a loop in which the platform generates examples, trains candidate configurations, measures their results, and adjusts subsequent experiments.
Adaption reports that AutoScientist outperformed training configurations built by its research staff by an average of 35%. Across eight verticals and datasets ranging from 5,000 to 100,000 rows, reported win rates rose from 48% to 64%. The experiments used architectures available for fine-tuning through the company’s Together AI partnership.
Where prompt-first data fits
Invent a Dataset moves the starting point from an existing corpus to a behavioral specification. That workflow can reduce the cost of early experiments when the desired capability is clear but examples remain unavailable.
- Prototype a specialized capability before contracting annotation vendors.
- Create training data for proprietary workflows whose signals remain buried in unlabeled internal logs.
- Generate preference pairs before a product has enough traffic to supply human comparisons.
- Produce locale-aware examples for markets with limited native-language data.
Hosted generation still needs scrutiny
Adaption documents generation as a hosted, credit-based service, with maximum row counts determined by plan tier. No self-hosted deployment path is documented. Teams with privacy, residency, or procurement constraints should verify the platform’s data-handling terms before submitting sensitive behavior specifications.
Generated rows also require task-specific validation. Developers should inspect label accuracy, factual errors, unsafe content, distribution gaps, and performance on held-out evaluations before using the data in production training. Regulated and safety-critical applications may require expert review alongside automated checks.
Teams whose main problem is cleaning and labeling existing records may still benefit more from conventional data pipelines. Teams with a precise target behavior and no seed dataset gain a direct route from specification to downloadable training rows, with generation, localization, and training optimization available through the same platform.