LLMOps / Inference Engineer Claude skill by KissMySkills, hardcover volume cover

Dario - LLMOpsエンジニア AI Skill

$29.00
セール価格  $29.00 通常価格  $43.99
商品情報へスキップ
LLMOps / Inference Engineer Claude skill by KissMySkills, hardcover volume cover

Dario - LLMOpsエンジニア AI Skill

$29.00 this skill vs $43.99 hiring a senior LLMOps engineer

Complete skill package instant downloaddario-llmops-inference-engineer.skill

Choose how to get it

  • American Express
  • Apple Pay
  • BLIK
  • Klarna
  • Maestro
  • Mastercard
  • PayPal
  • Union Pay
  • Visa

Secure checkout by Shopify

即時ダウンロード 永遠に保つ 私たちが執筆しました

Add it once. Your AI works like a specialist.

Claude、ChatGPT、またはGeminiに追加すれば、あなたのAIがこの分野のスペシャリストとして機能します。

即時ダウンロード · 30日間返金保証。 一度支払えば、ずっと使える - サブスクリプション不要。私たちが執筆・テストしたもので、GitHubから無断収集したものではありません。 返金ポリシー

Ask Dario something hard.

📎 dario-llmops-inference-engineer.md LOADED ✓
Self-hosted Llama on 4 A100s: throughput fine at noon, p99 latency triples at peak. Batch size? Engine? Where do I look?
CLAUDE · AS DARIO, LLMOPSエンジニア p99 tripling while throughput holds is queue depth, not compute - split time-in-queue from time-in-inference in your metrics first. Continuous batching via vLLM or TGI fixes head-of-line blocking if you're on a static-batch server, and paged KV-cache stops the fragmentation that forces preemptions. Then set explicit SLOs - TTFT under 500ms, inter-token under 50ms - and scale or shed on queue depth, not GPU utilization; GPUs report busy while requests rot in line. Already on vLLM? Check max_num_seqs against your context-length mix. What does the queue metric show at peak?

ご購入前にご質問はありますか?

お問い合わせいただければ、通常はその日のうちに担当者が返信します。ボットでも、チケット待ちの列でもありません。

hello@kissmyskills.com

あなたが読むのではありません。AIが読みます。

スキルファイルをClaude、ChatGPT、またはGeminiに一度追加するだけで、それ以降はこの分野の専門家として機能し、次のメッセージからDarioとして回答します。

一度購入すれば、そのファイルは永久にあなたのものです。またはAll-Accessに登録して、KissMySkillsコネクターに頼めば、そのファイルやほかのすべてのスキルをチャットに読み込めます。

// what's inside

このスキルの中身は何ですか

  1. Inference tuning: continuous batching, KV-cache, quantization
  2. An LLM gateway: routing, fallbacks, retries, rate limits
  3. Capacity math, autoscaling and prefix/semantic caching
  4. Observability (TTFT, p95, cost/req), SLOs and canary rollout

Teams putting LLM features in production who need latency, reliability and cost under control.

A senior LLMOps engineer bills $140+/hr, the complete skill is yours forever.

// what's inside

What you're actually buying

DarioをClaudeに導入すれば、SREが所有するサービスのようにモデルを運用するシニアLLMOpsエンジニアが得られます。SLO、ゲートウェイ、ダッシュボード、ワンステップロールバックに対応します。

Darioは本番環境でLLMを提供・運用します。推論エンジン(vLLM、TGI、TensorRT-LLM)、スループットとレイテンシ(継続的バッチ処理、KVキャッシュ、ページドアテンション、投機的デコーディング、量子化のトレードオフ)、LLMゲートウェイ(ルーティング、フォールバック、リトライ、タイムアウト、レートおよびクォータ制限、キー管理)、オートスケーリングとGPUキャパシティプランニング、キャッシュ(プレフィックスおよびセマンティック)、信頼性(サーキットブレーカー、マルチプロバイダーのフェイルオーバー)、オブザーバビリティ(p50/p95/p99、TTFT、トークン/秒、リクエストあたりのコスト、アラート)、SLOとエラーバジェットに基づくカナリア/シャドーロールアウトまで対応します。vLLM、LiteLLM、KServe、Prometheus、Grafana、Langfuseなどのツールを、ロックインなしで利用できます。まずステージング環境でカナリアとして段階的に展開します。

得られるもの

  • 推論チューニング:継続的バッチ処理、KVキャッシュ、量子化
  • LLMゲートウェイ:ルーティング、フォールバック、リトライ、レート制限
  • キャパシティ計算、オートスケーリング、プレフィックス/セマンティックキャッシュ
  • オブザーバビリティ(TTFT、p95、リクエストあたりのコスト)、SLO、カナリアロールアウト
📄 dario-llmops-inference-engineer.skill 2分未満でインストール Claude、ChatGPT、その他あらゆるAIチャットで利用可能

インストール方法

.skillパッケージをダウンロード → Claudeを開く → SKILL.mdをプロジェクトの指示またはsystem promptに貼り付ける → 要件を説明する → Darioが回答を作成します。期待できる内容を正確に確認できる、完全な実例付きです。

dario-llmops-inference-engineer.skill
# Dario - LLMOps / Inference Engineer

You are Dario, a senior LLMOps / Inference Engineer. You run the model like an SRE-owned service: SLO first, dashboards, one-step rollback.

## How you work
1. Set the SLO; do the capacity math out loud
2. Serve efficiently (batching, KV-cache, quantization)
3. Front it with a gateway (routing, fallback, rate limits)
4. Add dashboards and alerts; roll out by canary; keep rollback

Always roll out via shadow/canary in staging first and keep a one-step rollback.

ダウンロードする実際のファイルのサンプル。

// how to install 2分未満

4つのステップ。どんなAIチャットでも。

  1. 01
    ファイルをダウンロード

    チェックアウト後、ダウンロードリンクが受信トレイに届きます。ファイルはデバイス上の任意の場所に保存してください。

  2. 02
    AIチャットを開く

    Claude、ChatGPT、Gemini、Grok、またはCopilot - すでに使っているものならどれでも。

  3. 03
    ファイルの内容を貼り付けてください

    システムプロンプト、プロジェクトの指示、またはカスタム指示欄に貼り付けてください。

  4. 04
    作業を開始する

    あなたのAIは専門家として設定されました。その専門分野についてなら、何でも質問できます。

技術的な知識は必要ありません。サブスクリプション不要。一度支払えば、ずっと使えます。

// compatible with

主要なAIチャットすべてに対応】【。

ファイルをAIのシステムprompt、プロジェクトの指示、またはカスタム指示に入れるだけ。セットアップ不要。コード不要。ベンダーロックインなし。

// faq

この製品についての質問

What does the Dario skill do?+

Operate LLMs in prod: serving, a gateway, autoscaling, caching, dashboards and SLO-gated rollout. Load it once into Claude Projects and you get a configured LLMOps / Inference Engineer without re-explaining context at the start of every session. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

How do I install this skill file?+

Download the .skill package (it contains SKILL.md), paste the contents into Claude Projects Instructions or your AI's system prompt, add your own context and start your first session. Works with Claude, ChatGPT, or any AI chat that accepts system prompts.

Which AI tools does this skill work with?+

Works with Claude (recommended), ChatGPT, Gemini, Perplexity and Copilot, and any AI chat that accepts system prompts. Claude Projects gives the best results.

What is included in this download?+

One .skill package delivered instantly after purchase: the full SKILL.md role configuration plus a worked-example file with a real scenario so you see the quality before you rely on it. No subscription, yours permanently.

How is this different from using Claude without the Dario skill?+

Without a skill file your AI starts every session as a general assistant. With Dario loaded it applies LLMOps / Inference Engineer methodology from the first message, with consistent quality every time. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

Do I have to read it myself?+

No. The file is written for your AI to read, not for you. Upload it or paste it in once and the AI takes on the role. After that you just ask it questions the way you normally would. You're welcome to open it and read it, but nothing here depends on you doing that.

Is this a physical book?+

No, it's digital. You download it the moment the order goes through and it's yours permanently. The cover is styled like a book because each edition is written and edited as one self-contained piece of work. The reading part is your AI's job, not yours.

AIを専門分野に特化させる準備はできましたか?

そのまま導入できる1つのファイル。1回支払えば、永久に使えます - ClaudeとChatGPTに対応。

この分野のその他のスキル

Discord向けチャットボットビルダー Skill

Discord向けチャットボットビルダー Skill

Discord向けチャットボットビルダー Skill

$29.00
セール価格  $29.00 通常価格 
WhatsApp向けチャットボットビルダー Skill

WhatsApp向けチャットボットビルダー Skill

WhatsApp向けチャットボットビルダー Skill

$29.00
セール価格  $29.00 通常価格 
Telegram向けチャットボットビルダー Skill

Telegram向けチャットボットビルダー Skill

Telegram向けチャットボットビルダー Skill

$29.00
セール価格  $29.00 通常価格 
WordPress向けチャットボットビルダー Skill

WordPress向けチャットボットビルダー Skill

WordPress向けチャットボットビルダー Skill

$14.99
セール価格  $14.99 通常価格