Ingrid — LLM Evaluation Engineer AI Skill
Instant download · 30-day money-back guarantee. Pay once, keep forever — no subscription. Refund policy
Evaluate LLM apps: golden sets, LLM-as-judge, RAG and agent metrics, and a CI gate that blocks regressions.
- Golden datasets: coverage, edge cases, no leakage
- LLM-as-judge: rubric, bias mitigation, human calibration
- RAG and agent metrics (faithfulness, tool-call accuracy)
- CI regression gates with confidence intervals and significance
Teams shipping prompt and model changes blind, who need a real eval gate before release.A senior LLM eval engineer bills $130+/hr, this is one file, yours forever.
Drop Ingrid into Claude and get a senior LLM evaluation engineer whose rule is simple: no ship without an eval gate.
Ingrid evaluates LLM apps and agents: building golden datasets (coverage, edge cases, no leakage, versioning), metric selection, LLM-as-judge (rubric design, pairwise vs pointwise, bias and position effects with mitigations, calibration against human labels), RAG metrics (faithfulness, context and answer relevance), agent trajectory eval (tool-call correctness, task completion), offline harnesses and CI regression gates, online eval (A/B, guardrail metrics, human review sampling), and statistical rigor (confidence intervals, significance). Tools like Ragas, promptfoo, DeepEval, LangSmith and Braintrust, without lock-in. Always evaluate on a held-out set before shipping.
What you get
- →Golden datasets: coverage, edge cases, no leakage
- →LLM-as-judge: rubric, bias mitigation, human calibration
- →RAG and agent metrics (faithfulness, tool-call accuracy)
- →CI regression gates with confidence intervals and significance
How to install
Download the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Ingrid builds the answer. Includes a full worked example so you see exactly what you get.
# Ingrid - LLM Evaluation Engineer You are Ingrid, a senior LLM Evaluation Engineer. Your rule: no ship without an eval gate. ## How you work 1. Build a golden set (coverage, edge cases, no leakage) 2. Pick metrics; design an LLM-as-judge rubric; calibrate to humans 3. Run offline; report numbers with confidence intervals 4. Gate CI on regressions; confirm online with A/B and sampling Always evaluate on a held-out set in dev/staging before shipping to production.
Excerpt from the actual file you'll download.
Four steps. Any AI chat.
- 01Download the file
After checkout, the download link lands in your inbox. Save the file anywhere on your device.
- 02Open your AI chat
Claude, ChatGPT, Gemini, Grok, or Copilot — whichever one you already use.
- 03Paste the file contents
Drop it into the system prompt, Project instructions, or custom instructions field.
- 04Start working
Your AI is now configured as a specialist. Ask it anything inside its domain.
No technical knowledge required. No subscription. Pay once, keep forever.
Works with every major AI chat.
Drop the file into your AI's system prompt, Project instructions, or custom instructions. No setup. No code. No vendor lock-in.
- Claude
- ChatGPT
- Gemini
- Grok
- Copilot
Works with any AI chat that accepts a system prompt or custom instructions.
Ready to specialise your AI?
One drop-in file. Pay once, keep forever — works with Claude & ChatGPT.