Real-time AI meeting assistant
A Windows app that listens to any call at the OS level — Teams, Zoom, Meet or a softphone, no bot in the meeting — and drafts a grounded answer about two seconds after the customer asks. Answers draw on company context and the exact text on screen, and stream in from Claude while the rep is still listening.
- 01ListenWASAPI loopback + mic, any meeting app, no bot
- 02TranscribeAzure AI Speech, speaker labelled Others / Me
- 03Groundscreen text via UI Automation + company context
- 04DraftClaude streams the answer to a card or mini bar
- 05Hand offtranscript and summary out to the CRM
The problem
On discovery calls, customers ask specific questions: “Does this integrate with our ERP?”, “What does the Enterprise tier cost?”, “How are you different from what we already have?” The rep either knows the answer, says “let me get back to you”, or searches product docs while the customer waits. After the call, notes typed into the CRM from memory lose the details that mattered.
So the job had three parts: a live transcript that knows who is speaking, on-demand answers grounded in the company’s own material, and clean output that goes straight into business systems.
What I built
- OS-level listening. WASAPI loopback captures the meeting app’s audio and labels it Others; the microphone is Me. The speaker comes from the audio source, not from guessing. It works with any meeting tool and needs no IT approval.
- One-hotkey answers. Five styles: bullets, paragraph, one-liner, story, or ask back, which suggests smart counter-questions. Auto mode spots new questions and drafts an answer without a key press.
- Screen grounding. Screen capture plus Windows UI Automation reads the exact text of a CRM record, spreadsheet or slide, so a number on a shared pricing sheet is quoted, not guessed.
- Context presets. Up to 30,000 characters of standing instructions (products, price rules, discount limits, competitor notes), plus per-meeting notes. One click switches presets.
- Phone mode. Azure Conversation Transcription with speaker diarization for PC calls and speakerphones.
- Out of the way. An always-on-top mini bar keeps the meeting window clear, and transcripts of up to 3,000 lines are ready to paste into the CRM or summarise.
Engineering for latency
An answer that arrives eight seconds after the question is a meeting summary, not an assistant.
- Streaming everywhere. Tokens render while the model is still writing. A lighter model handles live calls; a larger one handles document work.
- A short prompt. Only the last ~6,000 characters of transcript are sent, not the whole session. Prompt caching is on for multi-turn sessions and skipped for one-shot drafts.
- Audio normalisation. Meeting apps output float32 in many PCM formats, so everything is converted to 16 kHz mono 16-bit before speech recognition.
- End-of-turn timing. Loopback sends nothing during silence, so short silent blocks are injected to close phrases on time. Auto mode waits for ~1.8 s of quiet (5 s at most) and keeps at least 6 s between automatic answers.
- No freezes. The recognizer shuts down on the thread pool with a timeout.
Guardrails
- Screen reads are bounded. They are capped by time and size, password and input fields are never read, sensitive windows (banking, HR) are blocklisted, and the user sees the exact text before it is sent.
- Prompts live in one embedded JSON file. Code only picks a variant (language, style) and fills placeholders in one pass, so text pasted into the context panel cannot inject new instructions. A fixed answer format (question, topic, or “no question”) drives the card layout.
- Secrets are encrypted per Windows user with DPAPI. Enterprise deployments use their own Azure Speech resource and Claude API account.
- Humans approve every write. Reads can be automatic; writes to business systems need a confirmation preview.
What’s next
- Tool use through MCP servers: CRM lookups at call start and deal updates after; quotes from the real price book; follow-ups booked on the calendar while the call is live; action items turned into Jira, Asana or Linear tickets; answers from the company wiki with sources.
- An AI phone receptionist. The same audio → transcript → reasoning pipeline, plus telephony (SIP, Twilio, Azure Communication Services), natural text-to-speech and barge-in. It would answer opening hours and prices, take bookings and phone orders, and hand overflow calls to staff with a written summary, in English and Vietnamese.
Lessons
Live beats after-the-fact: the value is in the two seconds after a question, not in the summary an hour later. Listening at the OS level works with every tool without a bot. Grounding is the feature, because generic models give generic answers. And building the pipeline once means the same engine can power a sales copilot today and a phone receptionist tomorrow.