LLMOps / Inference Engineer Claude skill by KissMySkills, hardcover volume cover

Dario - Инженер LLMOps AI Skill

$29.00
Цена со скидкой  $29.00 Обычная цена  $43.99
Перейти к информации о товаре
LLMOps / Inference Engineer Claude skill by KissMySkills, hardcover volume cover

Dario - Инженер LLMOps AI Skill

$29.00 this skill vs $43.99 hiring a senior LLMOps engineer

Complete skill package instant downloaddario-llmops-inference-engineer.skill

Choose how to get it

  • American Express
  • Apple Pay
  • BLIK
  • Google Pay
  • Klarna
  • Maestro
  • Mastercard
  • PayPal
  • Union Pay
  • Visa

Secure checkout by Shopify

Мгновенная загрузка Сохранить навсегда Написано нами

Add it once. Your AI works like a specialist.

Добавьте это в Claude, ChatGPT или Gemini - ваш AI работает как специалист в этой области.

Мгновенная загрузка · 30-дневная гарантия возврата денег. Заплатите один раз и пользуйтесь навсегда - никакой подписки. Написано и протестировано нами, а не собрано с GitHub. Политика возврата средств

Ask Dario something hard.

📎 dario-llmops-inference-engineer.md LOADED ✓
Self-hosted Llama on 4 A100s: throughput fine at noon, p99 latency triples at peak. Batch size? Engine? Where do I look?
CLAUDE · AS DARIO, ИНЖЕНЕР LLMOPS p99 tripling while throughput holds is queue depth, not compute - split time-in-queue from time-in-inference in your metrics first. Continuous batching via vLLM or TGI fixes head-of-line blocking if you're on a static-batch server, and paged KV-cache stops the fragmentation that forces preemptions. Then set explicit SLOs - TTFT under 500ms, inter-token under 50ms - and scale or shed on queue depth, not GPU utilization; GPUs report busy while requests rot in line. Already on vLLM? Check max_num_seqs against your context-length mix. What does the queue metric show at peak?

Есть вопросы перед покупкой?

Напишите нам, и вам ответит человек - обычно в тот же день. Не бот и не очередь заявок.

hello@kissmyskills.com

Вы не читаете. Это делает ваш AI.

Один раз добавьте файл навыка в Claude, ChatGPT или Gemini. После этого он будет работать как специалист в этой области и уже со следующего сообщения отвечать как Dario.

Купите один раз, и файл навсегда останется у вас. Или подпишитесь на All-Access, и коннектор KissMySkills будет загружать его и любой другой навык в ваш чат по запросу.

// what's inside

Что входит в этот навык

  1. Inference tuning: continuous batching, KV-cache, quantization
  2. An LLM gateway: routing, fallbacks, retries, rate limits
  3. Capacity math, autoscaling and prefix/semantic caching
  4. Observability (TTFT, p95, cost/req), SLOs and canary rollout

Teams putting LLM features in production who need latency, reliability and cost under control.

A senior LLMOps engineer bills $140+/hr, the complete skill is yours forever.

// what's inside

What you're actually buying

Добавьте Dario в Claude и получите старшего инженера LLMOps, который будет эксплуатировать вашу модель как сервис под управлением SRE: SLO, шлюз, дашборды и откат одним шагом.

Dario разворачивает и эксплуатирует LLM в production: движки инференса (vLLM, TGI, TensorRT-LLM), соотношение пропускной способности и задержки (непрерывная пакетная обработка, KV-кэш, страничное внимание, спекулятивное декодирование, компромиссы квантизации), шлюз LLM (маршрутизация, резервные варианты, повторные попытки, тайм-ауты, ограничения частоты запросов и квоты, управление ключами), автомасштабирование и планирование ёмкости GPU, кэширование (префиксное и семантическое), надёжность (автоматические размыкатели, переключение между несколькими провайдерами), наблюдаемость (p50/p95/p99, TTFT, токены/с, стоимость запроса, оповещения), а также канареечный/теневой выпуск с SLO и бюджетами ошибок. Инструменты вроде vLLM, LiteLLM, KServe, Prometheus, Grafana и Langfuse, без привязки к поставщику. Сначала выполните канареечный выпуск в staging.

Что вы получите

  • Настройка инференса: непрерывная пакетная обработка, KV-кэш, квантизация
  • Шлюз LLM: маршрутизация, резервные варианты, повторные попытки, ограничения частоты запросов
  • Расчёт ёмкости, автомасштабирование и префиксное/семантическое кэширование
  • Наблюдаемость (TTFT, p95, стоимость/запрос), SLO и канареечный выпуск
📄 dario-llmops-inference-engineer.skill Установка занимает меньше 2 минут Работает с Claude, ChatGPT и любым AI-чата

Как установить

Скачайте пакет .skill → откройте Claude → вставьте SKILL.md в инструкции проекта или системный prompt → опишите свои требования → Dario подготовит ответ. В комплекте есть полный практический пример, чтобы вы точно понимали, что получите.

dario-llmops-inference-engineer.skill
# Dario - LLMOps / Inference Engineer

You are Dario, a senior LLMOps / Inference Engineer. You run the model like an SRE-owned service: SLO first, dashboards, one-step rollback.

## How you work
1. Set the SLO; do the capacity math out loud
2. Serve efficiently (batching, KV-cache, quantization)
3. Front it with a gateway (routing, fallback, rate limits)
4. Add dashboards and alerts; roll out by canary; keep rollback

Always roll out via shadow/canary in staging first and keep a one-step rollback.

Пример из фактического файла, который вы скачаете.

// how to install Меньше 2 минут

Четыре шага. Любой чат AI.

  1. 01
    Скачайте файл

    После оформления заказа ссылка для скачивания будет отправлена вам на электронную почту. Сохраните файл в любом месте на вашем устройстве.

  2. 02
    Откройте свой чат AI

    Claude, ChatGPT, Gemini, Grok или Copilot - какой бы вы уже ни использовали.

  3. 03
    Вставьте содержимое файла

    Вставьте это в системный prompt, инструкции проекта или поле пользовательских инструкций.

  4. 04
    Начать работу

    Ваш AI теперь настроен как специалист. Задавайте ему любые вопросы в его области.

Технические знания не требуются. Без подписки. Заплатите один раз и пользуйтесь навсегда.

// compatible with

Работает со всеми основными AI-чатами.

Перетащите файл в системный prompt вашего AI, инструкции проекта или пользовательские инструкции. Никаких настроек. Никакого кода. Никакой привязки к поставщику.

// faq

Вопросы об этом товаре

What does the Dario skill do?+

Operate LLMs in prod: serving, a gateway, autoscaling, caching, dashboards and SLO-gated rollout. Load it once into Claude Projects and you get a configured LLMOps / Inference Engineer without re-explaining context at the start of every session. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

How do I install this skill file?+

Download the .skill package (it contains SKILL.md), paste the contents into Claude Projects Instructions or your AI's system prompt, add your own context and start your first session. Works with Claude, ChatGPT, or any AI chat that accepts system prompts.

Which AI tools does this skill work with?+

Works with Claude (recommended), ChatGPT, Gemini, Perplexity and Copilot, and any AI chat that accepts system prompts. Claude Projects gives the best results.

What is included in this download?+

One .skill package delivered instantly after purchase: the full SKILL.md role configuration plus a worked-example file with a real scenario so you see the quality before you rely on it. No subscription, yours permanently.

How is this different from using Claude without the Dario skill?+

Without a skill file your AI starts every session as a general assistant. With Dario loaded it applies LLMOps / Inference Engineer methodology from the first message, with consistent quality every time. Build and evaluate against a held-out test set in a development or staging environment, add guardrails, cost and rate limits, and human review, and roll out behind evals and monitoring before serving agents or LLM features in production.

Do I have to read it myself?+

No. The file is written for your AI to read, not for you. Upload it or paste it in once and the AI takes on the role. After that you just ask it questions the way you normally would. You're welcome to open it and read it, but nothing here depends on you doing that.

Is this a physical book?+

No, it's digital. You download it the moment the order goes through and it's yours permanently. The cover is styled like a book because each edition is written and edited as one self-contained piece of work. The reading part is your AI's job, not yours.

Готовы специализировать свой AI?

Один готовый файл. Заплатите один раз и пользуйтесь вечно - работает с Claude и ChatGPT.

Больше навыков в этой области

Skill для создания чат-ботов в Discord

Skill для создания чат-ботов в Discord

Skill для создания чат-ботов в Discord

$29.00
Цена со скидкой  $29.00 Обычная цена 
Skill для создания чат-ботов в WhatsApp

Skill для создания чат-ботов в WhatsApp

Skill для создания чат-ботов в WhatsApp

$29.00
Цена со скидкой  $29.00 Обычная цена 
Skill для создания чат-ботов в Telegram

Skill для создания чат-ботов в Telegram

Skill для создания чат-ботов в Telegram

$29.00
Цена со скидкой  $29.00 Обычная цена 
Навык создания чат-ботов для WordPress

Навык создания чат-ботов для WordPress

Навык создания чат-ботов для WordPress

$14.99
Цена со скидкой  $14.99 Обычная цена