Thibaut — Eval Dataset & Synthetic Data Engineer AI Skill
Instant download · 30-day money-back guarantee. Pay once, keep forever — no subscription. Refund policy
Build trustworthy eval and training data: representative sampling, annotation quality, synthetic generation with real gates.
- Representative held-out sampling and failure-mode stratification
- Annotation guidelines and inter-annotator agreement
- Synthetic generation with diversity and label-noise gates
- Contamination checks, versioning and golden-set upkeep
Teams whose eval scores look great but do not predict production behavior at all.A senior data curation engineer bills $130+/hr, this is one file, yours forever.
Drop Thibaut into Claude and get a data engineer who builds the eval set your scores are only as trustworthy as.
Thibaut builds and curates the datasets underneath every eval and every fine-tune: sampling a held-out set that actually represents production traffic, stratifying by the failure modes you care about, annotation guidelines and inter-annotator agreement, synthetic data generation with diversity and label-noise gates, contamination and leakage checks between train and eval, dataset versioning and provenance, golden-set maintenance as the product changes, and knowing when a synthetic batch should be thrown away rather than shipped.
What you get
- →Representative held-out sampling and failure-mode stratification
- →Annotation guidelines and inter-annotator agreement
- →Synthetic generation with diversity and label-noise gates
- →Contamination checks, versioning and golden-set upkeep
How to install
Download the .skill package, open Claude, paste SKILL.md into your Project Instructions or system prompt, describe your requirement, and Thibaut builds the answer. Includes a full worked example so you see exactly what you get.
# Thibaut - Eval Dataset & Synthetic Data Engineer You are Thibaut, an Eval Dataset & Synthetic Data Engineer. A score on an unrepresentative eval set is worse than no score. ## How you work 1. Compare the eval set against real production traffic distribution 2. Stratify by the failure modes that actually matter 3. Write annotation guidelines; measure inter-annotator agreement 4. Generate synthetic data behind diversity and label-noise gates Check for train/eval contamination before trusting any improvement.
Excerpt from the actual file you'll download.
Four steps. Any AI chat.
- 01Download the file
After checkout, the download link lands in your inbox. Save the file anywhere on your device.
- 02Open your AI chat
Claude, ChatGPT, Gemini, Grok, or Copilot — whichever one you already use.
- 03Paste the file contents
Drop it into the system prompt, Project instructions, or custom instructions field.
- 04Start working
Your AI is now configured as a specialist. Ask it anything inside its domain.
No technical knowledge required. No subscription. Pay once, keep forever.
Works with every major AI chat.
Drop the file into your AI's system prompt, Project instructions, or custom instructions. No setup. No code. No vendor lock-in.
- Claude
- ChatGPT
- Gemini
- Grok
- Copilot
Works with any AI chat that accepts a system prompt or custom instructions.
Ready to specialise your AI?
One drop-in file. Pay once, keep forever — works with Claude & ChatGPT.