Xiulan — Speech & Audio ML Engineer AI Skill
Instant download · 30-day money-back guarantee. Pay once, keep forever — no subscription. Refund policy
Ship speech that works: tune ASR for your domain, measure WER by accent and condition, and hold the latency budget.
- ASR tuning: domain adaptation, custom vocabulary, biasing
- WER measured by accent, noise condition and speaker, not headline
- Diarization, speaker ID, VAD and channel separation
- Streaming pipelines, latency budgets, TTS and consent limits
Teams shipping voice or call-audio products whose transcription fails on the accents and noise their real users have.A speech and audio ML engineer commands $145+/hr, this is one file, yours forever.
Drop Xiulan into Claude and get a speech engineer who reports word error rate by accent and condition, because a single headline WER hides exactly the users you are failing.
Xiulan owns the audio modality end to end: automatic speech recognition with Whisper and its successors, streaming versus batch decoding, WER measurement and what it conceals, domain adaptation, custom vocabulary and contextual biasing; speaker diarization and speaker identification; voice activity detection; text to speech and voice cloning under consent constraints; audio preprocessing including resampling, noise reduction, echo cancellation and channel separation; keyword spotting and wake words; audio classification and event detection; real-time streaming pipelines with latency budgets; and the fairness problem of recognition quality varying sharply by accent and dialect.
What you get
- →ASR tuning: domain adaptation, custom vocabulary, biasing
- →WER measured by accent, noise condition and speaker, not headline
- →Diarization, speaker ID, VAD and channel separation
- →Streaming pipelines, latency budgets, TTS and consent limits
How to install
Download the .skill package → open Claude → paste SKILL.md into your Project Instructions or system prompt → describe your requirement → Xiulan builds the answer. Includes a full worked example so you see exactly what you get.
# Xiulan - Speech & Audio ML Engineer You are Xiulan, a speech and audio ML engineer. A headline WER is a marketing number; you always break it down by accent, noise and channel. ## How you work 1. Build a test set that matches your real callers, not clean read speech 2. Report WER sliced by accent, noise, channel and speaker 3. Fix preprocessing and channel separation before touching the model 4. Adapt the domain with custom vocabulary and biasing; measure the gain Never ship voice cloning or TTS of a real voice without explicit recorded consent.
Excerpt from the actual file you'll download.
Four steps. Any AI chat.
- 01Download the file
After checkout, the download link lands in your inbox. Save the file anywhere on your device.
- 02Open your AI chat
Claude, ChatGPT, Gemini, Grok, or Copilot — whichever one you already use.
- 03Paste the file contents
Drop it into the system prompt, Project instructions, or custom instructions field.
- 04Start working
Your AI is now configured as a specialist. Ask it anything inside its domain.
No technical knowledge required. No subscription. Pay once, keep forever.
Works with every major AI chat.
Drop the file into your AI's system prompt, Project instructions, or custom instructions. No setup. No code. No vendor lock-in.
- Claude
- ChatGPT
- Gemini
- Grok
- Copilot
Works with any AI chat that accepts a system prompt or custom instructions.
Ready to specialise your AI?
One drop-in file. Pay once, keep forever — works with Claude & ChatGPT.