LLMOps / Inference Engineer Claude skill by KissMySkills, hardcover volume cover

Dario - LLMOps 工程師 AI Skill

$29.00
特價  $29.00 原價  $43.99
跳至產品資訊
LLMOps / Inference Engineer Claude skill by KissMySkills, hardcover volume cover

Dario - LLMOps 工程師 AI Skill

$29.00 this skill vs $43.99 hiring a senior LLMOps engineer

Complete skill package instant downloaddario-llmops-inference-engineer.skill

Choose how to get it

  • American Express
  • Apple Pay
  • BLIK
  • Google Pay
  • Klarna
  • Maestro
  • Mastercard
  • PayPal
  • Union Pay
  • Visa

Secure checkout by Shopify

立即下載 永遠保留 由我們撰寫

Add it once. Your AI works like a specialist.

將其加入 Claude、ChatGPT 或 Gemini - 您的 AI 會在此領域中擔任專家。

立即下載 · 30 天退款保證。 一次付費,永久使用 - 無需訂閱。由我們撰寫並測試,絕非從 GitHub 抓取。 退款政策

Ask Dario something hard.

📎 dario-llmops-inference-engineer.md LOADED ✓
Self-hosted Llama on 4 A100s: throughput fine at noon, p99 latency triples at peak. Batch size? Engine? Where do I look?
CLAUDE · AS DARIO, LLMOPS 工程師 p99 tripling while throughput holds is queue depth, not compute - split time-in-queue from time-in-inference in your metrics first. Continuous batching via vLLM or TGI fixes head-of-line blocking if you're on a static-batch server, and paged KV-cache stops the fragmentation that forces preemptions. Then set explicit SLOs - TTFT under 500ms, inter-token under 50ms - and scale or shed on queue depth, not GPU utilization; GPUs report busy while requests rot in line. Already on vLLM? Check max_num_seqs against your context-length mix. What does the queue metric show at peak?

購買前有任何問題嗎?

寫信給我們,通常當天就會有人回覆。不會是機器人,也不是排隊等候的客服工單。

hello@kissmyskills.com

你不需要閱讀。由你的 AI 代勞。

只需將技能檔案新增至 Claude、ChatGPT 或 Gemini 一次。從那之後,它就會以這個領域的專家身分運作,並從下一則訊息開始以 Dario 的身分回答。

購買一次,檔案便永久歸您所有。或者訂閱 All-Access,讓 KissMySkills 連接器在您提出要求時,將該檔案及其他所有技能載入您的聊天中。

// what's inside

這項技能包含什麼

  1. Inference tuning: continuous batching, KV-cache, quantization
  2. An LLM gateway: routing, fallbacks, retries, rate limits
  3. Capacity math, autoscaling and prefix/semantic caching
  4. Observability (TTFT, p95, cost/req), SLOs and canary rollout

Teams putting LLM features in production who need latency, reliability and cost under control.

A senior LLMOps engineer bills $140+/hr, the complete skill is yours forever.

// what's inside

What you're actually buying

將 Dario 放入 Claude,即可獲得一位資深 LLMOps 工程師,像由 SRE 負責的服務一樣運作您的模型:SLO、閘道、儀表板,以及一步回滾。

Dario 在正式環境中提供並運作 LLM:推理引擎(vLLM、TGI、TensorRT-LLM)、吞吐量與延遲(持續批次處理、KV-cache、分頁注意力、推測解碼、量化取捨)、LLM 閘道(路由、備援、重試、逾時、速率與配額限制、金鑰管理)、自動擴展與 GPU 容量規劃、快取(前綴與語意)、可靠性(斷路器、多供應商故障轉移)、可觀測性(p50/p95/p99、TTFT、每秒 token 數、每次請求成本、警示),以及搭配 SLO 與錯誤預算的金絲雀/影子發布。使用 vLLM、LiteLLM、KServe、Prometheus、Grafana 與 Langfuse 等工具,不受供應商綁定。先在預備環境透過金絲雀發布推出。

您將獲得

  • 推理調校:持續批次處理、KV-cache、量化
  • LLM 閘道:路由、備援、重試、速率限制
  • 容量計算、自動擴展與前綴/語意快取
  • 可觀測性(TTFT、p95、每次請求成本)、SLO 與金絲雀發布
📄 dario-llmops-inference-engineer.skill 不到 2 分鐘即可安裝 適用於 Claude、ChatGPT 與任何 AI 聊天工具

如何安裝

下載 .skill 套件 → 開啟 Claude → 將 SKILL.md 貼到您的專案指示或 system prompt 中 → 描述您的需求 → Dario 建立答案。內含完整實作範例,讓您清楚了解實際成果。

dario-llmops-inference-engineer.skill
# Dario - LLMOps / Inference Engineer

You are Dario, a senior LLMOps / Inference Engineer. You run the model like an SRE-owned service: SLO first, dashboards, one-step rollback.

## How you work
1. Set the SLO; do the capacity math out loud
2. Serve efficiently (batching, KV-cache, quantization)
3. Front it with a gateway (routing, fallback, rate limits)
4. Add dashboards and alerts; roll out by canary; keep rollback

Always roll out via shadow/canary in staging first and keep a one-step rollback.

您將下載的實際檔案範例。

// how to install 不到 2 分鐘

四個步驟。任何 AI 聊天。

  1. 01
    下載檔案

    結帳後,下載連結會寄到您的收件匣。將檔案儲存至裝置上的任意位置。

  2. 02
    開啟你的 AI 聊天功能

    Claude、ChatGPT、Gemini、Grok 或 Copilot - 無論你目前使用哪一個。

  3. 03
    貼上檔案內容

    將其放入系統 prompt、專案指示或自訂指示欄位中。

  4. 04
    開始工作

    您的 AI 現已設定為專業助手。您可以詢問它任何與其專業領域相關的問題。

不需要技術知識。不需訂閱。一次付費,永久使用。

// compatible with

支援所有主要的 AI 聊天服務。

將檔案放入您的 AI system prompt、專案指示或自訂指示中。無需設定。無需程式碼。不受供應商綁定。

// faq

關於此產品的問題

What does the Dario skill do?+

Operate LLMs in prod: serving, a gateway, autoscaling, caching, dashboards and SLO-gated rollout. Load it once into Claude Projects and you get a configured LLMOps / Inference Engineer without re-explaining context at the start of every session. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

How do I install this skill file?+

Download the .skill package (it contains SKILL.md), paste the contents into Claude Projects Instructions or your AI's system prompt, add your own context and start your first session. Works with Claude, ChatGPT, or any AI chat that accepts system prompts.

Which AI tools does this skill work with?+

Works with Claude (recommended), ChatGPT, Gemini, Perplexity and Copilot, and any AI chat that accepts system prompts. Claude Projects gives the best results.

What is included in this download?+

One .skill package delivered instantly after purchase: the full SKILL.md role configuration plus a worked-example file with a real scenario so you see the quality before you rely on it. No subscription, yours permanently.

How is this different from using Claude without the Dario skill?+

Without a skill file your AI starts every session as a general assistant. With Dario loaded it applies LLMOps / Inference Engineer methodology from the first message, with consistent quality every time. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

Do I have to read it myself?+

No. The file is written for your AI to read, not for you. Upload it or paste it in once and the AI takes on the role. After that you just ask it questions the way you normally would. You're welcome to open it and read it, but nothing here depends on you doing that.

Is this a physical book?+

No, it's digital. You download it the moment the order goes through and it's yours permanently. The cover is styled like a book because each edition is written and edited as one self-contained piece of work. The reading part is your AI's job, not yours.

準備好讓您的 AI 專精了嗎?

一個即插即用的檔案。一次付費,永久保留 - 適用於 Claude 和 ChatGPT。

此領域的更多技能

Discord 聊天機器人建立 Skill

Discord 聊天機器人建立 Skill

Discord 聊天機器人建立 Skill

$29.00
特價  $29.00 原價 
WhatsApp 聊天機器人建立 Skill

WhatsApp 聊天機器人建立 Skill

WhatsApp 聊天機器人建立 Skill

$29.00
特價  $29.00 原價 
Telegram 聊天機器人建立 Skill

Telegram 聊天機器人建立 Skill

Telegram 聊天機器人建立 Skill

$29.00
特價  $29.00 原價 
適用於 WordPress 的聊天機器人建立 Skill

適用於 WordPress 的聊天機器人建立 Skill

適用於 WordPress 的聊天機器人建立 Skill

$14.99
特價  $14.99 原價