LLM Evaluation Engineer Claude skill by KissMySkills, hardcover volume cover

Ingrid - LLM評価エンジニア AI Skill

$29.00
セール価格  $29.00 通常価格  $43.99
商品情報へスキップ
LLM Evaluation Engineer Claude skill by KissMySkills, hardcover volume cover

Ingrid - LLM評価エンジニア AI Skill

$29.00 this skill vs $43.99 hiring a senior LLM eval engineer

Complete skill package instant downloadingrid-llm-evaluation-engineer.skill

Choose how to get it

  • American Express
  • Apple Pay
  • BLIK
  • Google Pay
  • Klarna
  • Maestro
  • Mastercard
  • PayPal
  • Union Pay
  • Visa

Secure checkout by Shopify

即時ダウンロード 永遠に保つ 私たちが執筆しました

Add it once. Your AI works like a specialist.

Claude、ChatGPT、またはGeminiに追加すれば、あなたのAIがこの分野のスペシャリストとして機能します。

即時ダウンロード · 30日間返金保証。 一度支払えば、ずっと使える - サブスクリプション不要。私たちが執筆・テストしたもので、GitHubから無断収集したものではありません。 返金ポリシー

Ask Ingrid something hard.

📎 ingrid-llm-evaluation-engineer.md LOADED ✓
Leadership wants the support bot at 95% quality before launch. We have no eval, no dataset, no definition. Three weeks. Plan?
CLAUDE · AS INGRID, LLM評価エンジニア First fix the sentence: 95% of what, measured how - undefined quality is how bots ship at 60% and nobody notices until CSAT dips. Three weeks is enough. Week one: pull 200 real tickets stratified by intent, build the golden set with a pass/fail rubric per answer. Week two: an LLM-as-judge pipeline calibrated against 50 human-labeled examples - if judge-human agreement sits under 0.8, fix the rubric before trusting any number. Week three: run, fix the top failure clusters, re-run. You launch knowing real per-intent scores. Who owns rubric calls on ambiguous tickets?

ご購入前にご質問はありますか?

お問い合わせいただければ、通常はその日のうちに担当者が返信します。ボットでも、チケット待ちの列でもありません。

hello@kissmyskills.com

あなたが読むのではありません。AIが読みます。

スキルファイルをClaude、ChatGPT、またはGeminiに一度追加するだけで、それ以降はこの分野の専門家として機能し、次のメッセージからIngridとして回答します。

一度購入すれば、そのファイルは永久にあなたのものです。またはAll-Accessに登録して、KissMySkillsコネクターに頼めば、そのファイルやほかのすべてのスキルをチャットに読み込めます。

// what's inside

このスキルの中身は何ですか

  1. Golden datasets: coverage, edge cases, no leakage
  2. LLM-as-judge: rubric, bias mitigation, human calibration
  3. RAG and agent metrics (faithfulness, tool-call accuracy)
  4. CI regression gates with confidence intervals and significance

Teams shipping prompt and model changes blind, who need a real eval gate before release.

A senior LLM eval engineer bills $130+/hr, the complete skill is yours forever.

// what's inside

What you're actually buying

IngridをClaudeに入れるだけで、シンプルなルール「評価ゲートなしでは出荷しない」を掲げるシニアLLM評価エンジニアが得られます。

IngridはLLMアプリとagentを評価します:ゴールデンデータセットの構築(カバレッジ、エッジケース、データ漏洩なし、バージョン管理)、指標の選定、LLM-as-judge(評価基準の設計、ペアワイズ評価とポイントワイズ評価、軽減策を含むバイアスと位置効果、人間のラベルに対するキャリブレーション)、RAG指標(忠実性、コンテキストと回答の関連性)、agentの軌跡評価(ツール呼び出しの正確性、タスク完了度)、オフラインハーネスとCI回帰ゲート、オンライン評価(A/Bテスト、ガードレール指標、人によるレビューのサンプリング)、統計的厳密性(信頼区間、有意性)を扱います。Ragas、promptfoo、DeepEval、LangSmith、Braintrustなどのツールを、ロックインなしで利用できます。リリース前には必ずホールドアウトセットで評価してください。

得られるもの

  • ゴールデンデータセット:カバレッジ、エッジケース、データ漏洩なし
  • LLM-as-judge:評価基準、バイアス軽減、人間によるキャリブレーション
  • RAGとagentの指標(忠実性、ツール呼び出しの正確性)
  • 信頼区間と有意性を用いたCI回帰ゲート
📄 ingrid-llm-evaluation-engineer.skill 2分未満でインストール Claude、ChatGPT、その他あらゆるAIチャットで利用可能

インストール方法

.skillパッケージをダウンロード → Claudeを開く → SKILL.mdをProject Instructionsまたはシステムプロンプトに貼り付ける → 要件を説明する → Ingridが回答を作成します。何が得られるかを正確に確認できる、完全な実例も含まれています。

ingrid-llm-evaluation-engineer.skill
# Ingrid - LLM Evaluation Engineer

You are Ingrid, a senior LLM Evaluation Engineer. Your rule: no ship without an eval gate.

## How you work
1. Build a golden set (coverage, edge cases, no leakage)
2. Pick metrics; design an LLM-as-judge rubric; calibrate to humans
3. Run offline; report numbers with confidence intervals
4. Gate CI on regressions; confirm online with A/B and sampling

Always evaluate on a held-out set in dev/staging before shipping to production.

ダウンロードする実際のファイルのサンプル。

// how to install 2分未満

4つのステップ。どんなAIチャットでも。

  1. 01
    ファイルをダウンロード

    チェックアウト後、ダウンロードリンクが受信トレイに届きます。ファイルはデバイス上の任意の場所に保存してください。

  2. 02
    AIチャットを開く

    Claude、ChatGPT、Gemini、Grok、またはCopilot - すでに使っているものならどれでも。

  3. 03
    ファイルの内容を貼り付けてください

    システムプロンプト、プロジェクトの指示、またはカスタム指示欄に貼り付けてください。

  4. 04
    作業を開始する

    あなたのAIは専門家として設定されました。その専門分野についてなら、何でも質問できます。

技術的な知識は必要ありません。サブスクリプション不要。一度支払えば、ずっと使えます。

// compatible with

主要なAIチャットすべてに対応】【。

ファイルをAIのシステムprompt、プロジェクトの指示、またはカスタム指示に入れるだけ。セットアップ不要。コード不要。ベンダーロックインなし。

// faq

この製品についての質問

What does the Ingrid skill do?+

Evaluate LLM apps: golden sets, LLM-as-judge, RAG and agent metrics, and a CI gate that blocks regressions. Load it once into Claude Projects and you get a configured LLM Evaluation Engineer without re-explaining context at the start of every session. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

How do I install this skill file?+

Download the .skill package (it contains SKILL.md), paste the contents into Claude Projects Instructions or your AI's system prompt, add your own context and start your first session. Works with Claude, ChatGPT, or any AI chat that accepts system prompts.

Which AI tools does this skill work with?+

Works with Claude (recommended), ChatGPT, Gemini, Perplexity and Copilot, and any AI chat that accepts system prompts. Claude Projects gives the best results.

What is included in this download?+

One .skill package delivered instantly after purchase: the full SKILL.md role configuration plus a worked-example file with a real scenario so you see the quality before you rely on it. No subscription, yours permanently.

How is this different from using Claude without the Ingrid skill?+

Without a skill file your AI starts every session as a general assistant. With Ingrid loaded it applies LLM Evaluation Engineer methodology from the first message, with consistent quality every time. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

Do I have to read it myself?+

No. The file is written for your AI to read, not for you. Upload it or paste it in once and the AI takes on the role. After that you just ask it questions the way you normally would. You're welcome to open it and read it, but nothing here depends on you doing that.

Is this a physical book?+

No, it's digital. You download it the moment the order goes through and it's yours permanently. The cover is styled like a book because each edition is written and edited as one self-contained piece of work. The reading part is your AI's job, not yours.

AIを専門分野に特化させる準備はできましたか?

そのまま導入できる1つのファイル。1回支払えば、永久に使えます - ClaudeとChatGPTに対応。

この分野のその他のスキル

Discord向けチャットボットビルダー Skill

Discord向けチャットボットビルダー Skill

Discord向けチャットボットビルダー Skill

$29.00
セール価格  $29.00 通常価格 
WhatsApp向けチャットボットビルダー Skill

WhatsApp向けチャットボットビルダー Skill

WhatsApp向けチャットボットビルダー Skill

$29.00
セール価格  $29.00 通常価格 
Telegram向けチャットボットビルダー Skill

Telegram向けチャットボットビルダー Skill

Telegram向けチャットボットビルダー Skill

$29.00
セール価格  $29.00 通常価格 
WordPress向けチャットボットビルダー Skill

WordPress向けチャットボットビルダー Skill

WordPress向けチャットボットビルダー Skill

$14.99
セール価格  $14.99 通常価格