Dario — LLMOps / Inference Engineer AI Skill
Instant download · 30-day money-back guarantee. Pay once, keep forever — no subscription. Refund policy
Operate LLMs in prod: serving, a gateway, autoscaling, caching, dashboards and SLO-gated rollout.
- Inference tuning: continuous batching, KV-cache, quantization
- An LLM gateway: routing, fallbacks, retries, rate limits
- Capacity math, autoscaling and prefix/semantic caching
- Observability (TTFT, p95, cost/req), SLOs and canary rollout
Teams putting LLM features in production who need latency, reliability and cost under control.A senior LLMOps engineer bills $140+/hr, this is one file, yours forever.
Drop Dario into Claude and get a senior LLMOps engineer who runs your model like an SRE-owned service: SLOs, a gateway, dashboards and one-step rollback.
Dario serves and operates LLMs in production: inference engines (vLLM, TGI, TensorRT-LLM), throughput vs latency (continuous batching, KV-cache, paged attention, speculative decoding, quantization tradeoffs), an LLM gateway (routing, fallbacks, retries, timeouts, rate and quota limits, key management), autoscaling and GPU capacity planning, caching (prefix and semantic), reliability (circuit breakers, multi-provider failover), observability (p50/p95/p99, TTFT, tokens/sec, cost per request, alerts), and canary/shadow rollout with SLOs and error budgets. Tools like vLLM, LiteLLM, KServe, Prometheus, Grafana and Langfuse, without lock-in. Roll out via canary in staging first.
What you get
- →Inference tuning: continuous batching, KV-cache, quantization
- →An LLM gateway: routing, fallbacks, retries, rate limits
- →Capacity math, autoscaling and prefix/semantic caching
- →Observability (TTFT, p95, cost/req), SLOs and canary rollout
How to install
Download the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Dario builds the answer. Includes a full worked example so you see exactly what you get.
# Dario - LLMOps / Inference Engineer You are Dario, a senior LLMOps / Inference Engineer. You run the model like an SRE-owned service: SLO first, dashboards, one-step rollback. ## How you work 1. Set the SLO; do the capacity math out loud 2. Serve efficiently (batching, KV-cache, quantization) 3. Front it with a gateway (routing, fallback, rate limits) 4. Add dashboards and alerts; roll out by canary; keep rollback Always roll out via shadow/canary in staging first and keep a one-step rollback.
Excerpt from the actual file you'll download.
Four steps. Any AI chat.
- 01Download the file
After checkout, the download link lands in your inbox. Save the file anywhere on your device.
- 02Open your AI chat
Claude, ChatGPT, Gemini, Grok, or Copilot — whichever one you already use.
- 03Paste the file contents
Drop it into the system prompt, Project instructions, or custom instructions field.
- 04Start working
Your AI is now configured as a specialist. Ask it anything inside its domain.
No technical knowledge required. No subscription. Pay once, keep forever.
Works with every major AI chat.
Drop the file into your AI's system prompt, Project instructions, or custom instructions. No setup. No code. No vendor lock-in.
- Claude
- ChatGPT
- Gemini
- Grok
- Copilot
Works with any AI chat that accepts a system prompt or custom instructions.
Ready to specialise your AI?
One drop-in file. Pay once, keep forever — works with Claude & ChatGPT.