TLTan LeBackend & automation · CalgaryLet’s talk
← All systemsSales & front desk · Windows desktop app

Real-time AI meeting assistant

A Windows app that listens to any call at the OS level — Teams, Zoom, Meet or a softphone, no bot in the meeting — and drafts a grounded answer about two seconds after the customer asks. Answers draw on company context and the exact text on screen, and stream in from Claude while the rep is still listening.

~2 sfrom question to draft answer
0bots joining the meeting
30kchars of company context per preset
Meeting assistant data flowMeeting audio and the microphone are captured with WASAPI, converted to 16 kHz mono PCM and transcribed by Azure AI Speech with speaker labels. The prompt builder combines the recent transcript, on-screen text read through UI Automation and the company context, then Claude streams an answer to the answer card, a call summary for the CRM, and in future MCP tool calls. The same pipeline will power an AI phone receptionist.~1.8 s of quiet ends the turnpreviewed before sendpasswords never readSOURCESMeeting audioTeams, Zoom, Meetspeaker: OthersMicrophoneyour own voicespeaker: MeScreen textUI AutomationCRM, sheets, slidesCompany contextproducts, pricingpresets, 30k charsCapture · NAudio WASAPIloopback + mic → 16 kHz mono PCMAzure AI Speechlive transcript with speaker labelsPrompt builderlast ~6,000 chars of transcript + screen text + contextanswer style · language · one embedded JSON templateClaude · streamingtokens render as they are writtenOUTPUTSAnswer card · mini bar~2 s after the questionbullets, story, ask-backCall summary3,000-line transcriptpasted into the CRMNext: MCP toolsCRM, calendar, ticketswrites need a confirmSame pipeline, next product: AI phone receptionist+ telephony (SIP, Twilio, ACS) · text-to-speech · barge-in · confirm before booking
  1. 01
    ListenWASAPI loopback + mic, any meeting app, no bot
  2. 02
    TranscribeAzure AI Speech, speaker labelled Others / Me
  3. 03
    Groundscreen text via UI Automation + company context
  4. 04
    DraftClaude streams the answer to a card or mini bar
  5. 05
    Hand offtranscript and summary out to the CRM

The problem

On discovery calls, customers ask specific questions: “Does this integrate with our ERP?”, “What does the Enterprise tier cost?”, “How are you different from what we already have?” The rep either knows the answer, says “let me get back to you”, or searches product docs while the customer waits. After the call, notes typed into the CRM from memory lose the details that mattered.

So the job had three parts: a live transcript that knows who is speaking, on-demand answers grounded in the company’s own material, and clean output that goes straight into business systems.

What I built

  • OS-level listening. WASAPI loopback captures the meeting app’s audio and labels it Others; the microphone is Me. The speaker comes from the audio source, not from guessing. It works with any meeting tool and needs no IT approval.
  • One-hotkey answers. Five styles: bullets, paragraph, one-liner, story, or ask back, which suggests smart counter-questions. Auto mode spots new questions and drafts an answer without a key press.
  • Screen grounding. Screen capture plus Windows UI Automation reads the exact text of a CRM record, spreadsheet or slide, so a number on a shared pricing sheet is quoted, not guessed.
  • Context presets. Up to 30,000 characters of standing instructions (products, price rules, discount limits, competitor notes), plus per-meeting notes. One click switches presets.
  • Phone mode. Azure Conversation Transcription with speaker diarization for PC calls and speakerphones.
  • Out of the way. An always-on-top mini bar keeps the meeting window clear, and transcripts of up to 3,000 lines are ready to paste into the CRM or summarise.

Engineering for latency

An answer that arrives eight seconds after the question is a meeting summary, not an assistant.

  • Streaming everywhere. Tokens render while the model is still writing. A lighter model handles live calls; a larger one handles document work.
  • A short prompt. Only the last ~6,000 characters of transcript are sent, not the whole session. Prompt caching is on for multi-turn sessions and skipped for one-shot drafts.
  • Audio normalisation. Meeting apps output float32 in many PCM formats, so everything is converted to 16 kHz mono 16-bit before speech recognition.
  • End-of-turn timing. Loopback sends nothing during silence, so short silent blocks are injected to close phrases on time. Auto mode waits for ~1.8 s of quiet (5 s at most) and keeps at least 6 s between automatic answers.
  • No freezes. The recognizer shuts down on the thread pool with a timeout.

Guardrails

  • Screen reads are bounded. They are capped by time and size, password and input fields are never read, sensitive windows (banking, HR) are blocklisted, and the user sees the exact text before it is sent.
  • Prompts live in one embedded JSON file. Code only picks a variant (language, style) and fills placeholders in one pass, so text pasted into the context panel cannot inject new instructions. A fixed answer format (question, topic, or “no question”) drives the card layout.
  • Secrets are encrypted per Windows user with DPAPI. Enterprise deployments use their own Azure Speech resource and Claude API account.
  • Humans approve every write. Reads can be automatic; writes to business systems need a confirmation preview.

What’s next

  • Tool use through MCP servers: CRM lookups at call start and deal updates after; quotes from the real price book; follow-ups booked on the calendar while the call is live; action items turned into Jira, Asana or Linear tickets; answers from the company wiki with sources.
  • An AI phone receptionist. The same audio → transcript → reasoning pipeline, plus telephony (SIP, Twilio, Azure Communication Services), natural text-to-speech and barge-in. It would answer opening hours and prices, take bookings and phone orders, and hand overflow calls to staff with a written summary, in English and Vietnamese.

Lessons

Live beats after-the-fact: the value is in the two seconds after a question, not in the summary an hour later. Listening at the OS level works with every tool without a bot. Grounding is the feature, because generic models give generic answers. And building the pipeline once means the same engine can power a sales copilot today and a phone receptionist tomorrow.

Contact · Let’s talk

Have a system that should run itself by now?

Tell me what people still do by hand every morning, or what breaks when nobody is watching. I'll map the shortest path to a job that just runs.

Project brief