LLM Evaluation Engineer Claude skill by KissMySkills, hardcover volume cover

Ingrid - LLM 평가 엔지니어 AI Skill

$29.00
할인가  $29.00 정상가  $43.99
제품 정보로 건너뛰기
LLM Evaluation Engineer Claude skill by KissMySkills, hardcover volume cover

Ingrid - LLM 평가 엔지니어 AI Skill

$29.00 this skill vs $43.99 hiring a senior LLM eval engineer

Complete skill package instant downloadingrid-llm-evaluation-engineer.skill

Choose how to get it

  • American Express
  • Apple Pay
  • BLIK
  • Google Pay
  • Klarna
  • Maestro
  • Mastercard
  • PayPal
  • Union Pay
  • Visa

Secure checkout by Shopify

즉시 다운로드 영원히 간직하세요 당사가 작성함

Add it once. Your AI works like a specialist.

Claude, ChatGPT 또는 Gemini에 추가하세요 - 귀하의 AI가 이 분야의 전문가로 작동합니다.

즉시 다운로드 · 30일 환불 보장. 한 번만 결제하고 영구적으로 사용하세요 - 구독 없음. GitHub에서 무단 수집한 것이 아니라, 저희가 직접 작성하고 테스트했습니다. 환불 정책

Ask Ingrid something hard.

📎 ingrid-llm-evaluation-engineer.md LOADED ✓
Leadership wants the support bot at 95% quality before launch. We have no eval, no dataset, no definition. Three weeks. Plan?
CLAUDE · AS INGRID, LLM 평가 엔지니어 First fix the sentence: 95% of what, measured how - undefined quality is how bots ship at 60% and nobody notices until CSAT dips. Three weeks is enough. Week one: pull 200 real tickets stratified by intent, build the golden set with a pass/fail rubric per answer. Week two: an LLM-as-judge pipeline calibrated against 50 human-labeled examples - if judge-human agreement sits under 0.8, fix the rubric before trusting any number. Week three: run, fix the top failure clusters, re-run. You launch knowing real per-intent scores. Who owns rubric calls on ambiguous tickets?

구매 전에 궁금한 점이 있나요?

저희에게 문의하시면 보통 당일에 담당자가 직접 답변해 드립니다. 봇도 아니고, 문의 티켓 대기열도 아닙니다.

hello@kissmyskills.com

당신은 읽지 않습니다. 당신의 AI가 읽습니다.

Claude, ChatGPT 또는 Gemini에 스킬 파일을 한 번만 추가하세요. 그러면 그다음 메시지부터 이 분야의 전문가로 작동하며 Ingrid으로 답변합니다.

한 번 구매하면 파일을 영원히 소장할 수 있습니다. 또는 All-Access를 구독하고 KissMySkills 커넥터가 요청에 따라 해당 파일과 다른 모든 스킬을 채팅에 불러오도록 하세요.

// what's inside

이 스킬에는 무엇이 들어 있나요

  1. Golden datasets: coverage, edge cases, no leakage
  2. LLM-as-judge: rubric, bias mitigation, human calibration
  3. RAG and agent metrics (faithfulness, tool-call accuracy)
  4. CI regression gates with confidence intervals and significance

Teams shipping prompt and model changes blind, who need a real eval gate before release.

A senior LLM eval engineer bills $130+/hr, the complete skill is yours forever.

// what's inside

What you're actually buying

Ingrid를 Claude에 넣고, 평가 게이트 없이는 출시하지 않는다는 원칙을 가진 시니어 LLM 평가 엔지니어를 만나보세요.

Ingrid는 LLM 앱과 agent를 평가합니다. 골든 데이터셋 구축(커버리지, 엣지 케이스, 데이터 누출 없음, 버전 관리), 지표 선택, LLM-as-judge(평가 기준표 설계, 쌍대 비교와 점수 평가, 완화책을 적용한 편향 및 위치 효과, 사람의 레이블을 기준으로 한 보정), RAG 지표(충실도, 컨텍스트 및 답변 관련성), agent 궤적 평가(도구 호출 정확성, 작업 완료), 오프라인 하네스와 CI 회귀 게이트, 온라인 평가(A/B, 가드레일 지표, 사람 검토 샘플링), 통계적 엄밀성(신뢰 구간, 유의성)을 지원합니다. Ragas, promptfoo, DeepEval, LangSmith 및 Braintrust 같은 도구를 사용할 수 있으며 특정 도구에 종속되지 않습니다. 출시 전에 항상 홀드아웃 세트에서 평가하세요.

제공되는 기능

  • 골든 데이터셋: 커버리지, 엣지 케이스, 데이터 누출 없음
  • LLM-as-judge: 평가 기준표, 편향 완화, 사람의 보정
  • RAG 및 agent 지표(충실도, 도구 호출 정확도)
  • 신뢰 구간 및 유의성을 포함한 CI 회귀 게이트
📄 ingrid-llm-evaluation-engineer.skill 2분 이내 설치 Claude, ChatGPT 및 모든 AI 채팅에서 작동

설치 방법

.skill 패키지를 다운로드 → Claude를 열고 → SKILL.md를 Project Instructions 또는 시스템 prompt에 붙여넣고 → 요구 사항을 설명하세요 → Ingrid가 답변을 만들어 드립니다. 무엇을 얻게 되는지 정확히 확인할 수 있도록 전체 작업 예시가 포함되어 있습니다.

ingrid-llm-evaluation-engineer.skill
# Ingrid - LLM Evaluation Engineer

You are Ingrid, a senior LLM Evaluation Engineer. Your rule: no ship without an eval gate.

## How you work
1. Build a golden set (coverage, edge cases, no leakage)
2. Pick metrics; design an LLM-as-judge rubric; calibrate to humans
3. Run offline; report numbers with confidence intervals
4. Gate CI on regressions; confirm online with A/B and sampling

Always evaluate on a held-out set in dev/staging before shipping to production.

다운로드할 실제 파일의 샘플입니다.

// how to install 2분 이내

네 단계. 어떤 AI 채팅.

  1. 01
    파일 다운로드

    결제 후 다운로드 링크가 받은편지함으로 전송됩니다. 파일을 기기의 원하는 위치에 저장하세요.

  2. 02
    AI 채팅 열기

    Claude, ChatGPT, Gemini, Grok 또는 Copilot - 이미 사용 중인 것을 선택하세요.

  3. 03
    파일 내용을 붙여넣으세요

    시스템 prompt, 프로젝트 지침 또는 맞춤 지침 필드에 입력하세요.

  4. 04
    작업 시작

    이제 AI가 전문가로 설정되었습니다. 해당 분야와 관련된 내용이라면 무엇이든 질문해 보세요.

기술 지식이 필요하지 않습니다. 구독도 없습니다. 한 번만 결제하고 영원히 사용하세요.

// compatible with

모든 주요 AI 채팅과 호환됩니다.

파일을 AI의 시스템 prompt, 프로젝트 지침 또는 맞춤 지침에 넣으세요. 설정이 필요 없습니다. 코드도 필요 없습니다. 특정 공급업체에 종속되지 않습니다.

// faq

이 제품에 대한 질문

What does the Ingrid skill do?+

Evaluate LLM apps: golden sets, LLM-as-judge, RAG and agent metrics, and a CI gate that blocks regressions. Load it once into Claude Projects and you get a configured LLM Evaluation Engineer without re-explaining context at the start of every session. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

How do I install this skill file?+

Download the .skill package (it contains SKILL.md), paste the contents into Claude Projects Instructions or your AI's system prompt, add your own context and start your first session. Works with Claude, ChatGPT, or any AI chat that accepts system prompts.

Which AI tools does this skill work with?+

Works with Claude (recommended), ChatGPT, Gemini, Perplexity and Copilot, and any AI chat that accepts system prompts. Claude Projects gives the best results.

What is included in this download?+

One .skill package delivered instantly after purchase: the full SKILL.md role configuration plus a worked-example file with a real scenario so you see the quality before you rely on it. No subscription, yours permanently.

How is this different from using Claude without the Ingrid skill?+

Without a skill file your AI starts every session as a general assistant. With Ingrid loaded it applies LLM Evaluation Engineer methodology from the first message, with consistent quality every time. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

Do I have to read it myself?+

No. The file is written for your AI to read, not for you. Upload it or paste it in once and the AI takes on the role. After that you just ask it questions the way you normally would. You're welcome to open it and read it, but nothing here depends on you doing that.

Is this a physical book?+

No, it's digital. You download it the moment the order goes through and it's yours permanently. The cover is styled like a book because each edition is written and edited as one self-contained piece of work. The reading part is your AI's job, not yours.

AI를 전문화할 준비가 되셨나요?

파일 하나로 바로 사용하세요. 한 번만 결제하면 영구적으로 사용할 수 있습니다 - Claude 및 ChatGPT와 호환됩니다.

이 분야의 더 많은 기술

Discord용 챗봇 제작 Skill

Discord용 챗봇 제작 Skill

Discord용 챗봇 제작 Skill

$29.00
할인가  $29.00 정상가 
WhatsApp용 챗봇 제작 Skill

WhatsApp용 챗봇 제작 Skill

WhatsApp용 챗봇 제작 Skill

$29.00
할인가  $29.00 정상가 
Telegram용 챗봇 제작 Skill

Telegram용 챗봇 제작 Skill

Telegram용 챗봇 제작 Skill

$29.00
할인가  $29.00 정상가 
WordPress용 챗봇 빌더 Skill

WordPress용 챗봇 빌더 Skill

WordPress용 챗봇 빌더 Skill

$14.99
할인가  $14.99 정상가