Bertil — Model Distillation & On-Device Engineer AI Skill
Instant download · 30-day money-back guarantee. Pay once, keep forever — no subscription. Refund policy
Get a small model shipping: distill it, quantize it with the quality cost measured, and hold the device budget.
- Response, feature and rationale distillation from a teacher
- INT8/INT4, GPTQ, AWQ and GGUF with measured quality cost
- llama.cpp, MLX, Core ML, ONNX and TFLite on real hardware
- Memory, thermal and battery budgets; hybrid escalation patterns
Teams that need a model running offline or on-device without paying for it in accuracy they did not measure.An on-device model engineer commands $165+/hr, this is one file, yours forever.
Drop Bertil into Claude and get a distillation engineer who measures the quality cost of every quantization step instead of assuming INT4 is free.
Bertil makes a small model good enough and gets it onto the device: knowledge distillation from a large teacher across response, feature and rationale distillation; task-specific small language models; quantization across INT8, INT4, GPTQ, AWQ and GGUF with the quality cost measured rather than assumed; pruning and sparsity; LoRA and adapter merging; on-device runtimes including llama.cpp, MLX, ONNX Runtime, Core ML and TFLite; memory, thermal and battery budgets on real hardware; latency on mobile and edge silicon; the hybrid pattern where a small model handles the common path and escalates the rest; offline capability and update distribution; and honest evaluation against the teacher on the tasks that matter.
What you get
- →Response, feature and rationale distillation from a teacher
- →INT8/INT4, GPTQ, AWQ and GGUF with measured quality cost
- →llama.cpp, MLX, Core ML, ONNX and TFLite on real hardware
- →Memory, thermal and battery budgets; hybrid escalation patterns
How to install
Download the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Bertil builds the answer. Includes a full worked example so you see exactly what you get.
# Bertil - Model Distillation & On-Device Engineer You are Bertil, a distillation and on-device engineer. Every compression step has a quality cost and you measure it rather than assuming it is free. ## How you work 1. Define the target tasks first; general benchmarks will mislead you 2. Distill from the teacher, then quantize step by step with measurement 3. Test on the real device: latency, memory, thermals, battery 4. Design the hybrid escalation path for what the small model cannot do Never ship a quantization level you have not evaluated on the target tasks, and state plainly which tasks the small model fails.
Excerpt from the actual file you'll download.
Four steps. Any AI chat.
- 01Download the file
After checkout, the download link lands in your inbox. Save the file anywhere on your device.
- 02Open your AI chat
Claude, ChatGPT, Gemini, Grok, or Copilot — whichever one you already use.
- 03Paste the file contents
Drop it into the system prompt, Project instructions, or custom instructions field.
- 04Start working
Your AI is now configured as a specialist. Ask it anything inside its domain.
No technical knowledge required. No subscription. Pay once, keep forever.
Works with every major AI chat.
Drop the file into your AI's system prompt, Project instructions, or custom instructions. No setup. No code. No vendor lock-in.
- Claude
- ChatGPT
- Gemini
- Grok
- Copilot
Works with any AI chat that accepts a system prompt or custom instructions.
Ready to specialise your AI?
One drop-in file. Pay once, keep forever — works with Claude & ChatGPT.