{"product_id":"ingrid-llm-evaluation-engineer","title":"Ingrid — LLM Evaluation Engineer AI Skill","description":"\u003cdiv style=\"font-family: 'DM Sans', sans-serif; color: #1A1A18; max-width: 680px;\"\u003e\n  \u003cp style=\"font-size: 16px; font-weight: 600; line-height: 1.5; margin: 0 0 8px 0;\"\u003eDrop Ingrid into Claude and get a senior LLM evaluation engineer whose rule is simple: no ship without an eval gate.\u003c\/p\u003e\n  \u003cp style=\"font-size: 13px; color: #555550; line-height: 1.7; margin: 0 0 28px 0;\"\u003eIngrid evaluates LLM apps and agents: building golden datasets (coverage, edge cases, no leakage, versioning), metric selection, LLM-as-judge (rubric design, pairwise vs pointwise, bias and position effects with mitigations, calibration against human labels), RAG metrics (faithfulness, context and answer relevance), agent trajectory eval (tool-call correctness, task completion), offline harnesses and CI regression gates, online eval (A\/B, guardrail metrics, human review sampling), and statistical rigor (confidence intervals, significance). Tools like Ragas, promptfoo, DeepEval, LangSmith and Braintrust, without lock-in. Always evaluate on a held-out set before shipping.\u003c\/p\u003e\n  \u003cdiv style=\"background: #ECEDFC; border-radius: 12px; padding: 24px 28px; margin-bottom: 24px;\"\u003e\n    \u003cp style=\"font-size: 10px; font-weight: 600; color: #4A5BEE; letter-spacing: 0.08em; text-transform: uppercase; margin: 0 0 16px 0;\"\u003eWhat you get\u003c\/p\u003e\n    \u003cul style=\"margin: 0; padding: 0; list-style: none;\"\u003e\n\u003cli style=\"font-size: 13px; padding: 7px 0; border-bottom: 1px solid rgba(74,91,238,0.14); display: flex; gap: 10px;\"\u003e\n\u003cspan style=\"color:#4A5BEE; font-weight:600;\"\u003e→\u003c\/span\u003e\u003cspan\u003eGolden datasets: coverage, edge cases, no leakage\u003c\/span\u003e\n\u003c\/li\u003e\n\u003cli style=\"font-size: 13px; padding: 7px 0; border-bottom: 1px solid rgba(74,91,238,0.14); display: flex; gap: 10px;\"\u003e\n\u003cspan style=\"color:#4A5BEE; font-weight:600;\"\u003e→\u003c\/span\u003e\u003cspan\u003eLLM-as-judge: rubric, bias mitigation, human calibration\u003c\/span\u003e\n\u003c\/li\u003e\n\u003cli style=\"font-size: 13px; padding: 7px 0; border-bottom: 1px solid rgba(74,91,238,0.14); display: flex; gap: 10px;\"\u003e\n\u003cspan style=\"color:#4A5BEE; font-weight:600;\"\u003e→\u003c\/span\u003e\u003cspan\u003eRAG and agent metrics (faithfulness, tool-call accuracy)\u003c\/span\u003e\n\u003c\/li\u003e\n\u003cli style=\"font-size: 13px; padding: 7px 0;  display: flex; gap: 10px;\"\u003e\n\u003cspan style=\"color:#4A5BEE; font-weight:600;\"\u003e→\u003c\/span\u003e\u003cspan\u003eCI regression gates with confidence intervals and significance\u003c\/span\u003e\n\u003c\/li\u003e\n    \u003c\/ul\u003e\n  \u003c\/div\u003e\n  \u003cdiv style=\"display:flex; align-items:center; gap:20px; background:#FFFFFF; border:1px solid #E8E6E0; border-radius:8px; padding:14px 20px; margin-bottom:24px;\"\u003e\n    \u003cspan style=\"font-size:11px; color:#888780; font-family:monospace;\"\u003e📄 ingrid-llm-evaluation-engineer.skill\u003c\/span\u003e\n    \u003cspan style=\"font-size:11px; color:#888780;\"\u003eUnder 2 min install\u003c\/span\u003e\n    \u003cspan style=\"font-size:11px; color:#888780;\"\u003eWorks with Claude, ChatGPT \u0026amp; any AI chat\u003c\/span\u003e\n  \u003c\/div\u003e\n  \u003cdiv style=\"border-left:3px solid #4A5BEE; padding-left:16px;\"\u003e\n    \u003cp style=\"font-size:10px; font-weight:600; color:#4A5BEE; letter-spacing:0.08em; text-transform:uppercase; margin:0 0 6px 0;\"\u003eHow to install\u003c\/p\u003e\n    \u003cp style=\"font-size:12px; color:#555550; line-height:1.7; margin:0;\"\u003eDownload the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Ingrid builds the answer. Includes a full worked example so you see exactly what you get.\u003c\/p\u003e\n  \u003c\/div\u003e\n\u003c\/div\u003e","brand":"KissMySkills","offers":[{"title":"Default Title","offer_id":58313000845576,"sku":null,"price":34.0,"currency_code":"USD","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/1036\/1444\/7880\/files\/ag2-ingrid-cover.png?v=1785164278","url":"https:\/\/kissmyskills.com\/zh\/products\/ingrid-llm-evaluation-engineer","provider":"KissMySkills","version":"1.0","type":"link"}