Ingrid - LLM 評估工程師 AI Skill
Complete skill package instant downloadingrid-llm-evaluation-engineer.skill
Choose how to get it
Start All-Access, from $8.25/mo$99 a year, or $19 a month. 14-day refund on the first charge.
// all-access
You do not download skills. Your AI fetches them.
The KissMySkills connector is a small server your AI chat talks to. Add it once to Claude, ChatGPT, Claude Code or Cursor. From then on, any of our 1,000+ skills is one sentence away.
- 01Add the connector. Settings, Connectors, paste one link. A minute.
- 02Sign in once. The email from your subscription. Nothing installs.
- 03Ask for a skill. "Load Ingrid." It arrives in the conversation and your AI answers as that specialist.
Every skill, prompt pack and agent, new releases included. Single files can still be bought and kept. Cancel any time from your account. Compare plans
Secure checkout by Shopify
Add it once. Your AI works like a specialist.
將其加入 Claude、ChatGPT 或 Gemini - 您的 AI 會在此領域中擔任專家。
立即下載 · 30 天退款保證。 一次付費,永久使用 - 無需訂閱。由我們撰寫並測試,絕非從 GitHub 抓取。 退款政策
Ask Ingrid something hard.
Trademarks of their respective owners. KissMySkills is not affiliated with or endorsed by them.
你不需要閱讀。由你的 AI 代勞。
只需將技能檔案新增至 Claude、ChatGPT 或 Gemini 一次。從那之後,它就會以這個領域的專家身分運作,並從下一則訊息開始以 Ingrid 的身分回答。
購買一次,檔案便永久歸您所有。或者訂閱 All-Access,讓 KissMySkills 連接器在您提出要求時,將該檔案及其他所有技能載入您的聊天中。
// what's inside
這項技能包含什麼
- Golden datasets: coverage, edge cases, no leakage
- LLM-as-judge: rubric, bias mitigation, human calibration
- RAG and agent metrics (faithfulness, tool-call accuracy)
- CI regression gates with confidence intervals and significance
Teams shipping prompt and model changes blind, who need a real eval gate before release.
A senior LLM eval engineer bills $130+/hr, the complete skill is yours forever.
What you're actually buying
將 Ingrid 放入 Claude,即可獲得一位資深 LLM 評估工程師,其原則很簡單:沒有評估閘門,就不上線。
Ingrid 評估 LLM 應用程式與 agent:建立黃金資料集(涵蓋範圍、邊界案例、無資料洩漏、版本控制)、指標選擇、LLM-as-judge(評分規準設計、成對比較與逐項評分、偏誤與位置效應及其緩解措施、與人工標註的校準)、RAG 指標(忠實度、上下文與答案相關性)、agent 軌跡評估(工具呼叫正確性、任務完成度)、離線測試框架與 CI 回歸閘門、線上評估(A/B、護欄指標、人工審查抽樣),以及統計嚴謹性(信賴區間、顯著性檢驗)。支援 Ragas、promptfoo、DeepEval、LangSmith 和 Braintrust 等工具,不受單一工具綁定。上線前務必在保留集上完成評估。
您將獲得
- →黃金資料集:涵蓋範圍、邊界案例、無資料洩漏
- →LLM-as-judge:評分規準、偏誤緩解、人工校準
- →RAG 與 agent 指標(忠實度、工具呼叫準確度)
- →具備信賴區間與顯著性檢驗的 CI 回歸閘門
如何安裝
下載 .skill 套件 → 開啟 Claude → 將 SKILL.md 貼到您的專案指示或系統 prompt 中 → 描述您的需求 → Ingrid 建構答案。包含完整的實作範例,讓您清楚了解實際成果。
# Ingrid - LLM Evaluation Engineer You are Ingrid, a senior LLM Evaluation Engineer. Your rule: no ship without an eval gate. ## How you work 1. Build a golden set (coverage, edge cases, no leakage) 2. Pick metrics; design an LLM-as-judge rubric; calibrate to humans 3. Run offline; report numbers with confidence intervals 4. Gate CI on regressions; confirm online with A/B and sampling Always evaluate on a held-out set in dev/staging before shipping to production.
您將下載的實際檔案範例。
四個步驟。任何 AI 聊天。
- 01下載檔案
結帳後,下載連結會寄到您的收件匣。將檔案儲存至裝置上的任意位置。
- 02開啟你的 AI 聊天功能
Claude、ChatGPT、Gemini、Grok 或 Copilot - 無論你目前使用哪一個。
- 03貼上檔案內容
將其放入系統 prompt、專案指示或自訂指示欄位中。
- 04開始工作
您的 AI 現已設定為專業助手。您可以詢問它任何與其專業領域相關的問題。
不需要技術知識。不需訂閱。一次付費,永久使用。
支援所有主要的 AI 聊天服務。
將檔案放入您的 AI system prompt、專案指示或自訂指示中。無需設定。無需程式碼。不受供應商綁定。
關於此產品的問題
What does the Ingrid skill do?+
Evaluate LLM apps: golden sets, LLM-as-judge, RAG and agent metrics, and a CI gate that blocks regressions. Load it once into Claude Projects and you get a configured LLM Evaluation Engineer without re-explaining context at the start of every session. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.
How do I install this skill file?+
Download the .skill package (it contains SKILL.md), paste the contents into Claude Projects Instructions or your AI's system prompt, add your own context and start your first session. Works with Claude, ChatGPT, or any AI chat that accepts system prompts.
Which AI tools does this skill work with?+
Works with Claude (recommended), ChatGPT, Gemini, Perplexity and Copilot, and any AI chat that accepts system prompts. Claude Projects gives the best results.
What is included in this download?+
One .skill package delivered instantly after purchase: the full SKILL.md role configuration plus a worked-example file with a real scenario so you see the quality before you rely on it. No subscription, yours permanently.
How is this different from using Claude without the Ingrid skill?+
Without a skill file your AI starts every session as a general assistant. With Ingrid loaded it applies LLM Evaluation Engineer methodology from the first message, with consistent quality every time. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.
Do I have to read it myself?+
No. The file is written for your AI to read, not for you. Upload it or paste it in once and the AI takes on the role. After that you just ask it questions the way you normally would. You're welcome to open it and read it, but nothing here depends on you doing that.
Is this a physical book?+
No, it's digital. You download it the moment the order goes through and it's yours permanently. The cover is styled like a book because each edition is written and edited as one self-contained piece of work. The reading part is your AI's job, not yours.
準備好讓您的 AI 專精了嗎?
一個即插即用的檔案。一次付費,永久保留 - 適用於 Claude 和 ChatGPT。