Ilinca — Multimodal Agent Engineer AI Skill
Instant download · 30-day money-back guarantee. Pay once, keep forever — no subscription. Refund policy
Ship multimodal agents: pick the right vision model, control token cost, and evaluate visual reasoning honestly.
- Vision-language model selection and real failure modes
- Resolution, tiling and token cost arithmetic per image
- Chart, table, screenshot and video understanding limits
- Multimodal retrieval, image hallucination patterns, honest evals
Product teams shipping image or document-aware agents that look impressive in a demo and fail on real inputs.A multimodal AI engineer bills $150+/hr, this is one file, yours forever.
Drop Ilinca into Claude and get a multimodal engineer who builds an eval set that tests visual reasoning, not caption matching, and knows where chart reading silently breaks.
Ilinca builds agents that see as well as read: vision-language model selection and their genuine failure modes; image preprocessing, resolution and tiling traded against token cost; chart, diagram and table understanding and exactly where it fails silently; screenshot and UI understanding for visual grounding; bounding boxes and the limits of spatial reasoning; video understanding through frame sampling strategy; combining OCR with a vision model rather than picking one; multimodal retrieval and image embeddings; evaluation sets that test reasoning rather than captioning; hallucination patterns specific to images; accessibility uses such as alt text; and the cost arithmetic when every image is thousands of tokens.
What you get
- →Vision-language model selection and real failure modes
- →Resolution, tiling and token cost arithmetic per image
- →Chart, table, screenshot and video understanding limits
- →Multimodal retrieval, image hallucination patterns, honest evals
How to install
Download the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Ilinca builds the answer. Includes a full worked example so you see exactly what you get.
# Ilinca - Multimodal Agent Engineer You are Ilinca, a multimodal agent engineer. You classify the failure before changing the model, and you price every image in tokens. ## How you work 1. Build an eval set that tests visual reasoning, not captioning 2. Separate OCR failure from spatial reasoning failure from hallucination 3. Tune resolution and tiling against measured token cost 4. Combine OCR with the vision model where the hybrid wins Never trust a vision model's reading of a chart without a validation rule, and always state cost per image before scaling.
Excerpt from the actual file you'll download.
Four steps. Any AI chat.
- 01Download the file
After checkout, the download link lands in your inbox. Save the file anywhere on your device.
- 02Open your AI chat
Claude, ChatGPT, Gemini, Grok, or Copilot — whichever one you already use.
- 03Paste the file contents
Drop it into the system prompt, Project instructions, or custom instructions field.
- 04Start working
Your AI is now configured as a specialist. Ask it anything inside its domain.
No technical knowledge required. No subscription. Pay once, keep forever.
Works with every major AI chat.
Drop the file into your AI's system prompt, Project instructions, or custom instructions. No setup. No code. No vendor lock-in.
- Claude
- ChatGPT
- Gemini
- Grok
- Copilot
Works with any AI chat that accepts a system prompt or custom instructions.
Ready to specialise your AI?
One drop-in file. Pay once, keep forever — works with Claude & ChatGPT.