VisitQuill: medical transcription research and product direction

Research date: September 7, 2026. Prices are USD unless stated otherwise. Working assumptions: U.S. practices, English first, iPhone capturing an in-person encounter. Remote consultations are a separate capture workflow. These assumptions await customer confirmation.

This is a review of current primary documentation, product pages, and published pricing. No vendors were contacted, subscriptions purchased, clinical recordings processed, or models benchmarked. Product capabilities below are vendor descriptions; recommendations and cost scenarios are our analysis. There is no demonstrated accuracy winner for this customer's audio yet.

Recommendation

Build an iOS encounter recorder with a verifiable, speaker-attributed transcript. Offer a separate dictation mode for the clinician's own notes and addenda. Make reliable capture, accurate attribution, and fast correction the first product. Preserve a structured transcript so EHR integration can follow without rebuilding the foundation.

Use native SwiftUI and AVFoundation for the first iPhone app. Evaluate Deepgram Nova-3 Medical and AssemblyAI Medical Mode as the initial cloud candidates, with Speechmatics as an additional accent/diarization comparator. Evaluate Apple SpeechAnalyzer plus a local diarizer, and FluidAudio, on actual iPhones alongside them. Keep providers interchangeable.

A cloud, local, or hybrid design must earn selection through measured clinical errors, speaker errors, recording reliability, and clinician correction time. Neither local processing nor low price is a reason to accept weaker clinical results. Evaluate latency, battery use, and total service cost after establishing the quality and healthcare deployment requirements. If a strictly local mode is selected, it must remain local even when it encounters difficulty.

VoiceStudio is excluded from the product direction. Do not install, benchmark, integrate, or pursue commercial licensing for it. Independently relevant speech engines remain eligible based on their fit for medical transcription; their inclusion in VoiceStudio gives them no preference.

1. What we are actually competing with

Three related products solve different jobs:

Product type Input and output VisitQuill scope
Dictation One clinician speaks; text appears or an existing note is edited Include as its own mode
Ambient transcription Natural conversation becomes a timestamped transcript with speakers Core first release
Ambient medical scribe Conversation and possibly chart context become structured clinical documentation Study the workflow now; generated notes and EHR writes come later

A human scribe may also review charts, recognize recurring patients, ask clarifying questions, use the doctor's preferred format, and complete administrative work. We have not verified which of these the current remote scribe service performs. Replacing its transcription is a smaller promise than replacing its entire service. Measure both the transcript and the doctor's remaining work before claiming superiority.

Products worth studying

This is a benchmark shortlist by useful capability, not an independently measured league table.

Product What its product teaches us Public price verified in this research
Abridge Clinical workflow depth and source verification. Its Linked Evidence feature connects note text to supporting transcript/audio. Build similarly fast verification into our transcript editor. Enterprise sales; no standard public dollar price found on reviewed pages. Product, Linked Evidence
Microsoft Dragon Copilot Combines ambient capture, clinician dictation, specialty-specific notes, and downstream tasks. Multi-party conversations and correction workflows matter alongside recognition accuracy. Contact sales; no standard public dollar price verified. Microsoft
Suki Voice-enabled editing, problem-based documentation, and clinician instructions. A doctor should be able to add information after the encounter without restarting it. Contact sales; no current standard public dollar price verified. Suki
Nabla Ambient documentation and dictation, customization, configurable retention, and an embeddable platform/API. Also a potential buy-versus-build partner. Quote required for the proposed embedded service; no current standard public subscription price verified. Nabla
Heidi Low-friction entry, templates, team workflows, and assistant seats. Free transcription establishes a demanding competitive baseline. Free tier; Clinician $150/user/month monthly, or $110/user/month billed yearly; Teams custom. USD selector and billing toggle verified in browser. U.S. pricing
Freed Simple onboarding, specialty templates, learning a preferred note format, dictation, and quick review. Strong small-practice price comparison. Starter $39/month, 40 notes; Core $79/month, unlimited notes; Premier $119/month monthly or $104/month annually; Groups custom. Browser verified; older hidden prices in the page's extracted text were disregarded. Pricing
DeepScribe Specialty-oriented documentation, customization, and workflows for fields such as oncology. Start with one specialty and its vocabulary rather than claiming every specialty immediately. Sales-led; no standard public dollar price verified. DeepScribe
Wispr Flow Fast activation and natural dictation are useful interaction references. Its current meeting Notetaker is Mac-only and explicitly not yet HIPAA-compliant; its dictation product has a separate HIPAA/BAA workflow. Growth dictation is $23/user/month monthly or $18 annually; dictation/Notetaker bundle $33 monthly or $26 annually. These are team Growth rates. Pricing, Notetaker FAQ

The Wispr distinction explains why a convenient desktop meeting experience cannot simply be treated as an approved medical capture workflow on iPhone. The desktop meeting workflow remains a useful interaction reference.

WHOIS / RDAP: domain registration history

All eight product domains were queried directly against their authoritative registries on September 7, 2026 (UTC). RDAP is the structured successor to WHOIS. The dates below use the registry's registration event, not the last update or expiration date. ICANN: About RDAP

Product / domain checked Registered (UTC; registry source) Domain age at research date How to interpret it
Abridge
abridge.com
1997-07-02 29 years, 2 months The registration date alone does not establish when the medical company acquired or began using this domain.
Microsoft Dragon Copilot
microsoft.com
1991-05-02 35 years, 4 months Parent-company domain. This date says nothing about the launch date of Dragon Copilot.
Suki
suki.ai
2017-12-16 8 years, 8 months Current brand domain; corroborate with company history and product deployments.
Nabla
nabla.com
1998-07-02 28 years, 2 months The registration date alone does not establish when the medical company acquired or began using this domain.
Heidi
heidihealth.com
2020-04-24 6 years, 4 months Current brand domain; earlier company names or domains would require separate research.
Freed
getfreed.ai
2023-01-16 3 years, 7 months Current brand domain; registration can precede product launch.
DeepScribe
deepscribe.ai
2019-04-27 7 years, 4 months Current brand domain; registration does not establish operating or clinical deployment history.
Wispr Flow
wisprflow.ai
2024-12-10 1 year, 8 months Current product domain; this young registration does not establish the age of the parent company.

Domain age is a limited maturity signal. These dates describe the current domain registration record. They do not prove company founding, product launch, continuous ownership, clinical experience, financial stability, or quality. Domains may be bought, transferred, repurposed, or re-registered after deletion. In particular, the 1990s registrations for Abridge and Nabla should not be presented as the ages of today's medical companies; Microsoft's corporate domain is also a poor proxy for Dragon Copilot's product age.

For vendor diligence, combine this history with verified launch dates, years of production deployments, customer references, published clinical evaluations, release history, and support commitments. Company founding dates and historical ownership chains have not been independently researched in this report.

Ages are completed calendar years and months as of September 7, 2026. Exact UTC timestamps and lookup times are preserved in the downloadable registry snapshot. This snapshot contains public registration metadata; registrant contact details are omitted. The linked registry responses are live and may change after the research date.

Features to match, and places to improve

Capability First-release behavior Later extension
Easy activation One clearly labeled start action; explicit active encounter; visible recording state Shortcuts, Action button, scheduling prompts
Speaker attribution Speaker A/B/C, confirmed clinician/patient/caregiver roles, timestamps Optional clinician enrollment; participant tracks in telehealth
Review Tap text to play matching audio; rename, merge, split, or reassign a speaker turn Evidence-linked clinical notes
Medical terminology Specialty vocabulary, drug names, abbreviations, numbers, units, and negation preservation Approved practice dictionaries and chart-derived hints
Reliable recording Lock-screen continuation, interruption alerts, encrypted recovery, offline queue Managed devices and room microphones
Clinician dictation Separate single-speaker mode; append a clearly identified addendum Voice commands and saved phrases
Uncertainty Mark unclear words, overlap, and unknown speakers; preserve original output Prioritized review using calibrated risk signals
User control Start/pause/stop, encounter boundaries, review state, retention choice within clinic policy Team review and delegated staff workflows
Portability Structured JSON and reviewed text export EHR adapters after patient matching and write review

Our proposed differentiator is lower verification effort with fewer attribution errors. This is a hypothesis to validate, not a claim that competitors lack these capabilities. Avoid making unsupported promises such as perfect accuracy or replacing all human scribe duties.

2. Speaker attribution is a system, not a checkbox

Transcription recognizes words. Diarization groups speech by speaker. Identification maps a voice to a person. Role assignment maps a participant to clinician, patient, caregiver, or interpreter. These outputs need separate fields and separate evaluation.

For example, a correct word sequence can still produce a dangerous record if a caregiver's statement about their own medication is attributed to the patient. An ASR model can also mistake "I do not take it" for "I do take it." General word accuracy does not adequately measure either problem.

Recommended behavior:

  1. Start with anonymous encounter-scoped speaker IDs. Confirm roles with the clinician; never assume the first speaker is the doctor.
  2. Show live labels as provisional when the engine can revise them. Stabilize the final transcript using the whole encounter where possible.
  3. Maintain explicit speaker mappings across chunks and reconnections. Provider "Speaker 0" from a restarted stream is not automatically the previous Speaker 0.
  4. Preserve overlap and unknown attribution. A forced confident label can be worse than a clear request for review.
  5. Keep acoustic confidence, speaker confidence, and clinical-review flags separate. A model's confidence is not a calibrated probability of clinical correctness.
  6. Let a reviewer correct one turn, a time interval, or a whole participant. Preserve changes and provenance.
  7. When a calling SDK supplies separate participant tracks, retain them. Channel separation generally gives stronger attribution evidence than inferring everyone from one mixed microphone, though a track can still contain more than one person.
  8. Consider optional clinician reference audio only after defining consent, storage, and deletion. Do not build persistent patient voiceprints for the MVP.

Pyannote's Community-1 documentation provides an offline diarization baseline and discusses an exclusive output that simplifies alignment with transcription. Exclusive output should not erase an overlap warning in our product. Model documentation

3. Speech APIs: capabilities and prices

Rates are public list/displayed rates observed on the research date, excluding taxes, free credits, minimum commitments, storage, retries, support, and negotiated healthcare terms. Per-hour conversions use 60 minutes. A cheap public API tier is not proof that our exact medical workflow is covered by a signed BAA.

Engine / configuration Published rate and normalized cost Fit and limitation
Deepgram Nova-3, batch $0.0043/min = $0.258/hour monolingual; batch diarization included Initial final-transcript candidate. Nova-3 Medical is a documented model option; its launch page quotes the same starting batch price. Confirm selected medical model, feature, and contract rates. Pricing, Medical model
Deepgram Nova-3, live + diarization Promotional $0.0048 + $0.0020 = $0.0068/min = $0.408/hour. Regular base displayed: $0.0077/min, making $0.0097/min with speakers Keyterm prompting adds $0.0013/min per pass. Do not budget as though the promotional rate is permanent. Pricing
AssemblyAI Universal-3.5 Pro, batch + speakers + Medical Mode $0.21 + $0.02 + $0.15 = $0.38/hour ≈ $0.00633/min Medical transcription comparator. Optional keyterms cost another $0.05/hour. Pricing
AssemblyAI Universal-3.5 Pro Realtime + speakers + Medical Mode $0.45 + $0.12 + $0.15 = $0.72/hour = $0.012/min Live medical/speaker candidate; keyterms included at this tier. Disable primary-speaker isolation for room conversations unless testing establishes it preserves all participants. Pricing
Speechmatics Displayed: Batch Melia 1 $0.129/hour; Standard batch/live $0.24/hour; Enhanced batch $0.40/hour, live $0.43/hour Evaluate accents and live speaker separation. Rates/features vary by model. Pricing has a model-training discount control; obtain a no-training healthcare quote. Enterprise deployment options include private infrastructure. Pricing
ElevenLabs Scribe v2 Batch $0.22/hour ≈ $0.00367/min; keyterms +$0.05/hour, entity detection +$0.07/hour Batch model documents speaker labels and word timestamps. A strong low-cost comparison candidate, subject to Scribe-specific BAA coverage. Pricing, Capabilities
ElevenLabs Scribe v2 Realtime $0.39/hour = $0.0065/min The documented realtime feature list does not establish parity with batch diarization. Do not promise live speaker attribution from the batch feature list. Same contract gate. Capabilities
OpenAI gpt-transcribe $0.0045/min = $0.27/hour Current recommended file model, with keyword/context hints; not a complete diarization solution. File guide, Pricing
OpenAI gpt-live-transcribe $0.017/min = $1.02/hour Current live model explicitly lacks speaker labels, word timestamps, and confidence scores. Would require a separate attribution/alignment solution. Realtime guide, Pricing
OpenAI gpt-4o-transcribe-diarize Estimated $0.006/min = $0.36/hour Existing file diarization endpoint, with optional known-speaker references. Deprecated; not a durable default for a new product. File guide
Amazon Transcribe Medical Published examples imply $0.075/min = $4.50/hour Medical dictation/conversation benchmark. Region, language, and model restrictions need checking for the actual deployment. Pricing examples
AWS HealthScribe $0.001667/second ≈ $0.10/min = $6/hour, 15-second minimum Bundles transcription, roles, entities, and evidence-linked clinical summaries. Broader and costlier than our current scope; documented batch and streaming, U.S. English, N. Virginia. Pricing, Technical scope
Google Cloud Speech-to-Text V2 / Chirp 3 Standard public starting rate $0.016/min = $0.96/hour Multilingual/diarization comparator; verify streaming feature availability for the selected language/region. Channels can be billed separately. Legacy medical models are distinct from Chirp 3. Overview and rate, Chirp 3, Billing
Azure Speech Unresolved here: the regional pricing page returned placeholders, not dollar values Viable streaming/batch and custom speech alternative, especially in Microsoft environments. The page lists realtime diarization as an add-on and batch diarization as included. Obtain a region-specific quote; do not confuse Azure Speech with the Dragon Copilot end-user product. Pricing

Two purchasing details that materially affect the shortlist

OpenAI migration: On August 26, 2026, OpenAI announced that whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize will be removed on February 26, 2027. The replacement transcription models do not establish speaker-feature parity. Keep an independent diarization route if evaluating OpenAI. This hosted API deprecation is separate from the availability of open Whisper weights. Deprecation notice

AssemblyAI BAA access: Its current Medical Mode documentation says upgraded paid accounts can sign a standard BAA from Data Controls at no additional cost. Older support/legal pages direct customers to sales. Verify availability and exact endpoint coverage in the proposed account. This could make AssemblyAI commercially easier to pilot than a vendor requiring an enterprise minimum. Medical Mode supports English, Spanish, German, and French; unsupported languages fall back with a warning. Medical Mode documentation, BAA page

Deepgram says to contact it for a BAA. ElevenLabs' cited HIPAA documentation concerns Agents, enterprise subscriptions, and zero-retention conditions; do not infer it automatically covers every standalone Scribe endpoint. OpenAI likewise requires checking the agreement and endpoint data controls, not treating an API key as healthcare authorization. Deepgram, ElevenLabs, OpenAI

4. Local model options

Local can mean three different things

Location Benefit Tradeoff
On the iPhone Works without uploading encounter audio; can work offline after models are downloaded Battery, heat, memory, model storage, background execution, and device support must be measured
A clinic-owned Mac or server Larger models with local operational control; phone can stream over the clinic network Pairing, network availability, updates, backups, and support become our responsibility
Our private server/cloud GPU Shared capacity and controlled model versions; portability Still a hosted PHI processor; infrastructure, security, and idle capacity are ongoing costs

Practical local candidates

Component Role Commercial cost and caveat
Apple SpeechAnalyzer / SpeechTranscriber On-device streaming and file ASR introduced in iOS 26; designed for longer conversations No metered ASR service bill for on-device use. Models must be downloaded; check device/language availability. Its reviewed public interface does not supply our complete diarization workflow. Apple WWDC
FluidAudio Swift/Core ML ASR, voice activity detection, speaker embeddings, and streaming/offline diarization; Parakeet-family ASR Apache-2.0 SDK; individual model terms also apply. Especially relevant to the iPhone prototype. Desktop speed claims are not iPhone battery or clinical-quality evidence. Repository
Argmax WhisperKit + SpeakerKit On-device Whisper transcription and Pyannote diarization for Apple platforms MIT open-source SDK plus third-party terms. Pro SDK advertises additional live features and Android support; commercial pricing needs inquiry. Repository
faster-whisper + WhisperX + Pyannote Self-hosted transcription, alignment, and speaker clustering Open components have no hosted per-minute tariff; budget hardware and operations. Review each code/model license. Better initial fit for a server or clinic Mac evaluation than direct iPhone embedding. faster-whisper, WhisperX, Pyannote

For local attribution, FluidAudio documents an offline Community-1 pipeline and streaming alternatives including LS-EEND and Sortformer. Select based on measured speaker stability and overlap behavior; a vendor's general benchmark ranking is not sufficient. Diarization implementation

VoiceStudio disposition

Excluded from the product direction. The initial review established a desktop voice-production workspace, without evidence that it is the best foundation for this medical iPhone workflow. No installation, evaluation, licensing work, or product dependency is planned. Repository reviewed during initial research

5. iPhone capture: the constraint to resolve first

Scenario Proposed approach
Doctor and patient in the same room Record the iPhone microphone in VisitQuill; test positioning and quiet speakers
Doctor's post-visit dictation Use a dedicated single-speaker session or labeled encounter addendum
Zoom/Teams running on another device The phone can hear room speakers, but headphones remove that acoustic path. Prefer a supported platform integration when remote visits become scope
Phone/FaceTime/Zoom/Teams call on the same iPhone Do not promise a generic cross-app audio tap. Use an explicitly supported meeting/telehealth integration, a call routed through infrastructure we control, or user-imported recordings where appropriate
Clinician walks through an eight-hour day Keep the app ready, but use explicit patient/encounter boundaries and pauses. One continuous recording risks mixing patients and increases review and processing costs

Apple documents that background audio recording can continue with the appropriate audio background mode, including when the screen locks, but calls and other nonmixable sessions can interrupt it. Our engineering conclusion is that uninterrupted microphone capture alongside an unrelated phone call is not a sound product assumption. This requires physical-device testing, not just a simulator demo. AVAudioSession recording behavior

If the remote human scribe is simply listening to an in-person visit through a phone left connected all day, VisitQuill can replace that capture arrangement by recording locally in the app. If the physician is actually consulting patients through that same phone call, the telephony path becomes a first-phase dependency. Confirm this distinction with the customer.

Apple also requires clear recording indication and user consent. A microphone permission prompt is not itself the full patient consent workflow. App Review Guidelines, recording/privacy

6. Architecture and future platforms

flowchart TD
  A[iPhone: explicit encounter and consent] --> B[AVFoundation recording and encrypted chunk journal]
  B --> C[Optional on-device live ASR and diarization]
  B --> D[Resumable upload to authenticated service]
  D --> E[Selected contracted ASR and diarization provider]
  C --> F[Versioned transcript and participant mapping]
  E --> F
  F --> G[Speaker correction, uncertainty review, audio playback]
  G --> H[Clinician-approved transcript and controlled export]
  H --> I[Later: clinical note generation and EHR adapter]

This is a proposal. In local-only mode, the cloud branch is disabled and cannot activate as a silent fallback.

Layer Initial choice Reason
iOS UI and audio SwiftUI, AVAudioSession/AVAudioEngine, Swift concurrency Direct control of the highest-risk part: recording, interruptions, routes, and native inference
Local persistence Encrypted audio chunks, protected metadata, Keychain-managed secrets Survive connectivity loss and process termination without losing encounter state
Service API Authenticated REST for records, WebSocket where needed for live updates Shared contract for later Android, Mac, Windows, and web clients
Processing Provider adapters and queued finalization jobs; Python workers are a practical initial option Switch engines without rewriting the app or transcript model
Records and audio Tenant-isolated relational database; encrypted object storage with lifecycle rules Separate access, retention, and revision history for audio and text
Operations Access logs, billing meters, model/version metadata, PHI-free operational telemetry Investigate failures and understand actual per-practice costs

SwiftUI is an iOS-first recommendation, not a promise of shared Android UI. Flutter or React Native could share more interface code, but background audio and local inference still require native work and device testing. Keep most reusable value in the service contract, transcript schema, review rules, and evaluation corpus. Android can later use Kotlin/native audio or a shared UI framework. Mac can reuse some Swift modules; Windows can use the same service and a separate capture adapter.

Preserve encounter IDs, provisional patient references, participant roles, timed segments, original model output, human edits, language, consent, and review state. A final EHR integration must independently confirm the patient, destination encounter, and clinician-approved content. We are deliberately building the data boundary now, not implementing chart writes.

7. What it costs to serve a doctor

Illustrative workload: 20 visits/day × 15 minutes × 22 days/month = 6,600 minutes = 110 audio hours/month. This is a planning assumption, not measured customer usage. Each full cloud pass processes the entire workload; running live and final models is two passes.

Configuration Speech processing / doctor / month 15-minute visit
Deepgram batch $28.38 $0.0645
Deepgram live with speakers, promotional $44.88 $0.1020
Deepgram live + complete batch final pass, promotional $73.26 $0.1665
Same two passes at displayed regular live rate $92.40 $0.2100
AssemblyAI medical batch with speakers $41.80 $0.0950
AssemblyAI medical live with speakers $79.20 $0.1800
AssemblyAI medical live + complete medical batch pass $121.00 $0.2750
ElevenLabs Scribe v2 batch $24.20 $0.0550
Speechmatics Enhanced live, displayed base rate $47.30 $0.1075
OpenAI gpt-transcribe, without separate diarization $29.70 $0.0675
OpenAI gpt-live-transcribe, without separate diarization $112.20 $0.2550
Google V2 standard starting rate $105.60 $0.2400
Amazon Transcribe Medical $495.00 $1.1250
AWS HealthScribe including its additional outputs Approximately $660.00 Approximately $1.50

Calculations use the cited rates in section 3; they are not vendor quotes. Optional vocabulary features, separate channels, medical contract minimums, and repeated retries can change these amounts. These configurations do not all deliver equivalent outputs or accuracy.

An eight-hour continuous daily stream for 22 days is 176 hours, or 1.6× this scenario, before extra passes. Do not assume voice-activity detection eliminates all billable silence: metering rules depend on the service and connection behavior.

Additional cost categories:

For a small pilot, reserve an illustrative $200–$600/month total shared infrastructure budget, excluding speech APIs, labor, vendor minimums, and legal/security work. This is our planning allowance, not measured infrastructure usage or a cloud quote. Allocate it across actual active clinicians and replace the allowance with a configured bill of materials before pricing a contract.

Local inference avoids a hosted per-minute recognition bill but still has model delivery, device support, QA, and engineering costs. For self-hosting, use:

effective cost/audio hour = (hardware or server cost + operations) / successfully processed audio hours

For illustration only, a server assumed to cost $1/hour and kept running 720 hours costs $720/month before operations. Compared with $0.258/audio-hour cloud batch recognition, hardware-only break-even is about 2,791 processed audio hours/month. Whether that server can meet peak concurrency and quality is unmeasured. Cheap cloud ASR can be hard to beat financially at small scale; local processing is initially strongest as a privacy/offline benefit.

Service pricing hypothesis

Test willingness to pay around $149–$249/clinician/month, with explicit included hours and usage-based team plans. This is a proposed discovery range, not a recommendation to launch at those prices. Freed and Heidi already offer clinical documentation at competitive prices; transcription alone needs compelling speaker-review quality or local privacy to justify a premium.

Do not promise unlimited all-day processing before measuring actual costs. At $199 revenue, $73.26 of speech processing and an assumed $40 of other service costs leaves roughly 43% contribution margin, before engineering and sales. A local live preview plus one cloud finalization pass could improve this, if it passes quality and device tests.

For distribution, direct organizational sales and individual subscriptions can have different App Store purchasing requirements. Apple's enterprise-services exception is specific to sales to organizations/groups; merely describing a subscription as a service does not create a universal exemption. Choose the sales channel before implementing billing. App Review Guidelines §3.1.3

8. Healthcare requirements that affect implementation now

U.S. medical audio can contain PHI before any EHR integration exists. HHS states that a cloud provider processing or storing ePHI can be a business associate even if it holds only encrypted data without the key. Signed agreements and appropriate technical/organizational controls belong in the transcription phase. Local inference reduces disclosure but does not alone establish compliance. HHS cloud guidance

Design the pilot around clinic-approved recording consent, tenant isolation, encryption, role-based access, clinician authentication, configurable retention, deletion of replicas and vendor copies, auditability, and an incident process. Keep PHI out of analytics, crash payloads, notifications, filenames, and routine support tooling. Review device backup and clipboard behavior. Retain raw audio only for the clinic-approved verification period; do not invent one retention duration for every practice.

The recorded transcript remains a draft until reviewed. Preserve original recognized text when presenting normalized numbers or suggested terminology corrections. Do not let a general-purpose LLM silently rewrite medications, doses, negations, or speaker identities. Keep diagnostic advice, autonomous coding, orders, and chart writes outside this first scope.

9. How we establish that it is better

Run a paired evaluation on the same audio, not separate demos with different microphones or speakers. Begin with role-play audio; move to consented clinical audio only within the agreed processing environment. The current human scribe is a useful comparator but is not automatically the error-free reference.

Use a held-out set covering the initial specialty, accented English, quiet/distant speech, masks, room noise, caregiver/interpreter participation, overlapping speech, negation, dosages, units, medication changes, and clinician addenda. Have qualified reviewers independently annotate the reference and resolve disagreements. Keep tuning examples separate from the final test set.

Measure:

Initial study proposal: 50–100 varied encounters for discovery, with a separately held-out evaluation portion. A small pilot can identify failure modes; it cannot establish universal clinical safety. Set specialty-specific acceptance thresholds with the clinical lead before evaluating vendors. Treat an unflagged critical speaker/medication error as a release-blocking finding to investigate, not something to hide inside average WER.

Suggested product performance targets, unmeasured and negotiable: visible live text within two seconds under defined network/device conditions; final transcript within one minute of stopping a 15-minute visit; every captured audio gap explicitly reported; no silent loss in the agreed recovery tests; materially less clinician correction time than the existing workflow. Quality gates take priority over these latency targets.

10. First build sequence

  1. Customer workflow and technical spike: confirm capture scenario, specialty, devices, language mix, and current scribe duties/cost. Build a native recorder with interruption/recovery instrumentation and run cloud/local comparisons on role-play recordings.
  2. Transcript MVP: encounter boundaries, participant mapping, recorded audio, live/final states, speaker editor, replay, addenda, and controlled export. Preserve provider-neutral records.
  3. Clinic pilot readiness: executed data agreements, approved consent/retention, access controls, deletion and recovery verification, support process, and clinician-led quality gates.
  4. Paired customer pilot: compare clinical and attribution errors, correction effort, and costs; select the production provider/mode and pricing from evidence.
  5. EHR phase: confirm patient identity and integrate reviewed content. Clinical note generation can be a distinct, evaluated stage.

Planning estimate: a focused technical spike may take 1–2 weeks, followed by several weeks of MVP work and pilot hardening for a small experienced team. Contracting, clinical evaluation, and device failures can dominate the calendar; this is not a delivery commitment.

See the accompanying transcription MVP blueprint for the concrete first-build acceptance criteria.

# VisitQuill: medical transcription research and product direction

Research date: **September 7, 2026**. Prices are USD unless stated otherwise. Working assumptions: U.S. practices, English first, iPhone capturing an in-person encounter. Remote consultations are a separate capture workflow. These assumptions await customer confirmation.

This is a review of current primary documentation, product pages, and published pricing. No vendors were contacted, subscriptions purchased, clinical recordings processed, or models benchmarked. Product capabilities below are vendor descriptions; recommendations and cost scenarios are our analysis. There is no demonstrated accuracy winner for this customer's audio yet.

## Recommendation

Build an **iOS encounter recorder with a verifiable, speaker-attributed transcript**. Offer a separate dictation mode for the clinician's own notes and addenda. Make reliable capture, accurate attribution, and fast correction the first product. Preserve a structured transcript so EHR integration can follow without rebuilding the foundation.

Use native SwiftUI and AVFoundation for the first iPhone app. Evaluate **Deepgram Nova-3 Medical and AssemblyAI Medical Mode** as the initial cloud candidates, with **Speechmatics** as an additional accent/diarization comparator. Evaluate **Apple SpeechAnalyzer plus a local diarizer**, and **FluidAudio**, on actual iPhones alongside them. Keep providers interchangeable.

A cloud, local, or hybrid design must earn selection through measured clinical errors, speaker errors, recording reliability, and clinician correction time. Neither local processing nor low price is a reason to accept weaker clinical results. Evaluate latency, battery use, and total service cost after establishing the quality and healthcare deployment requirements. If a strictly local mode is selected, it must remain local even when it encounters difficulty.

**VoiceStudio is excluded from the product direction.** Do not install, benchmark, integrate, or pursue commercial licensing for it. Independently relevant speech engines remain eligible based on their fit for medical transcription; their inclusion in VoiceStudio gives them no preference.

## 1. What we are actually competing with

Three related products solve different jobs:

| Product type | Input and output | VisitQuill scope |
| --- | --- | --- |
| Dictation | One clinician speaks; text appears or an existing note is edited | Include as its own mode |
| Ambient transcription | Natural conversation becomes a timestamped transcript with speakers | Core first release |
| Ambient medical scribe | Conversation and possibly chart context become structured clinical documentation | Study the workflow now; generated notes and EHR writes come later |

A human scribe may also review charts, recognize recurring patients, ask clarifying questions, use the doctor's preferred format, and complete administrative work. We have not verified which of these the current remote scribe service performs. Replacing its transcription is a smaller promise than replacing its entire service. Measure both the transcript and the doctor's remaining work before claiming superiority.

### Products worth studying

This is a benchmark shortlist by useful capability, not an independently measured league table.

| Product | What its product teaches us | Public price verified in this research |
| --- | --- | --- |
| **Abridge** | Clinical workflow depth and source verification. Its Linked Evidence feature connects note text to supporting transcript/audio. Build similarly fast verification into our transcript editor. | Enterprise sales; no standard public dollar price found on reviewed pages. [Product](https://www.abridge.com/platform/clinicians), [Linked Evidence](https://support.abridge.com/hc/en-us/articles/30235128433811-Verify-a-Note-With-Linked-Evidence) |
| **Microsoft Dragon Copilot** | Combines ambient capture, clinician dictation, specialty-specific notes, and downstream tasks. Multi-party conversations and correction workflows matter alongside recognition accuracy. | Contact sales; no standard public dollar price verified. [Microsoft](https://www.microsoft.com/en-us/health-solutions/clinical-workflow/dragon-copilot) |
| **Suki** | Voice-enabled editing, problem-based documentation, and clinician instructions. A doctor should be able to add information after the encounter without restarting it. | Contact sales; no current standard public dollar price verified. [Suki](https://www.suki.ai/) |
| **Nabla** | Ambient documentation and dictation, customization, configurable retention, and an embeddable platform/API. Also a potential buy-versus-build partner. | Quote required for the proposed embedded service; no current standard public subscription price verified. [Nabla](https://www.nabla.com/) |
| **Heidi** | Low-friction entry, templates, team workflows, and assistant seats. Free transcription establishes a demanding competitive baseline. | Free tier; Clinician **$150/user/month monthly**, or **$110/user/month billed yearly**; Teams custom. USD selector and billing toggle verified in browser. [U.S. pricing](https://www.heidihealth.com/en-us/pricing) |
| **Freed** | Simple onboarding, specialty templates, learning a preferred note format, dictation, and quick review. Strong small-practice price comparison. | Starter **$39/month**, 40 notes; Core **$79/month**, unlimited notes; Premier **$119/month monthly** or **$104/month annually**; Groups custom. Browser verified; older hidden prices in the page's extracted text were disregarded. [Pricing](https://www.getfreed.ai/pricing) |
| **DeepScribe** | Specialty-oriented documentation, customization, and workflows for fields such as oncology. Start with one specialty and its vocabulary rather than claiming every specialty immediately. | Sales-led; no standard public dollar price verified. [DeepScribe](https://www.deepscribe.ai/) |
| **Wispr Flow** | Fast activation and natural dictation are useful interaction references. Its current meeting Notetaker is Mac-only and explicitly not yet HIPAA-compliant; its dictation product has a separate HIPAA/BAA workflow. | Growth dictation is **$23/user/month monthly** or **$18 annually**; dictation/Notetaker bundle **$33 monthly** or **$26 annually**. These are team Growth rates. [Pricing](https://wisprflow.ai/enterprise-pricing), [Notetaker FAQ](https://docs.wisprflow.ai/articles/6858284702-notetaker-for-enterprise-faq) |

The Wispr distinction explains why a convenient desktop meeting experience cannot simply be treated as an approved medical capture workflow on iPhone. The desktop meeting workflow remains a useful interaction reference.

### WHOIS / RDAP: domain registration history

All eight product domains were queried directly against their authoritative registries on **September 7, 2026 (UTC)**. RDAP is the structured successor to WHOIS. The dates below use the registry's **registration** event, not the last update or expiration date. [ICANN: About RDAP](https://www.icann.org/rdap/)

| Product / domain checked | Registered (UTC; registry source) | Domain age at research date | How to interpret it |
| --- | --- | --- | --- |
| **Abridge**<br>abridge.com | [1997-07-02](https://rdap.verisign.com/com/v1/domain/abridge.com) | 29 years, 2 months | The registration date alone does not establish when the medical company acquired or began using this domain. |
| **Microsoft Dragon Copilot**<br>microsoft.com | [1991-05-02](https://rdap.verisign.com/com/v1/domain/microsoft.com) | 35 years, 4 months | Parent-company domain. This date says nothing about the launch date of Dragon Copilot. |
| **Suki**<br>suki.ai | [2017-12-16](https://rdap.identitydigital.services/rdap/domain/suki.ai) | 8 years, 8 months | Current brand domain; corroborate with company history and product deployments. |
| **Nabla**<br>nabla.com | [1998-07-02](https://rdap.verisign.com/com/v1/domain/nabla.com) | 28 years, 2 months | The registration date alone does not establish when the medical company acquired or began using this domain. |
| **Heidi**<br>heidihealth.com | [2020-04-24](https://rdap.verisign.com/com/v1/domain/heidihealth.com) | 6 years, 4 months | Current brand domain; earlier company names or domains would require separate research. |
| **Freed**<br>getfreed.ai | [2023-01-16](https://rdap.identitydigital.services/rdap/domain/getfreed.ai) | 3 years, 7 months | Current brand domain; registration can precede product launch. |
| **DeepScribe**<br>deepscribe.ai | [2019-04-27](https://rdap.identitydigital.services/rdap/domain/deepscribe.ai) | 7 years, 4 months | Current brand domain; registration does not establish operating or clinical deployment history. |
| **Wispr Flow**<br>wisprflow.ai | [2024-12-10](https://rdap.identitydigital.services/rdap/domain/wisprflow.ai) | 1 year, 8 months | Current product domain; this young registration does not establish the age of the parent company. |

**Domain age is a limited maturity signal.** These dates describe the current domain registration record. They do not prove company founding, product launch, continuous ownership, clinical experience, financial stability, or quality. Domains may be bought, transferred, repurposed, or re-registered after deletion. In particular, the 1990s registrations for Abridge and Nabla should not be presented as the ages of today's medical companies; Microsoft's corporate domain is also a poor proxy for Dragon Copilot's product age.

For vendor diligence, combine this history with verified launch dates, years of production deployments, customer references, published clinical evaluations, release history, and support commitments. Company founding dates and historical ownership chains have not been independently researched in this report.

Ages are completed calendar years and months as of September 7, 2026. Exact UTC timestamps and lookup times are preserved in the [downloadable registry snapshot](./domain-registration/domains.json). This snapshot contains public registration metadata; registrant contact details are omitted. The linked registry responses are live and may change after the research date.

### Features to match, and places to improve

| Capability | First-release behavior | Later extension |
| --- | --- | --- |
| Easy activation | One clearly labeled start action; explicit active encounter; visible recording state | Shortcuts, Action button, scheduling prompts |
| Speaker attribution | Speaker A/B/C, confirmed clinician/patient/caregiver roles, timestamps | Optional clinician enrollment; participant tracks in telehealth |
| Review | Tap text to play matching audio; rename, merge, split, or reassign a speaker turn | Evidence-linked clinical notes |
| Medical terminology | Specialty vocabulary, drug names, abbreviations, numbers, units, and negation preservation | Approved practice dictionaries and chart-derived hints |
| Reliable recording | Lock-screen continuation, interruption alerts, encrypted recovery, offline queue | Managed devices and room microphones |
| Clinician dictation | Separate single-speaker mode; append a clearly identified addendum | Voice commands and saved phrases |
| Uncertainty | Mark unclear words, overlap, and unknown speakers; preserve original output | Prioritized review using calibrated risk signals |
| User control | Start/pause/stop, encounter boundaries, review state, retention choice within clinic policy | Team review and delegated staff workflows |
| Portability | Structured JSON and reviewed text export | EHR adapters after patient matching and write review |

Our proposed differentiator is **lower verification effort with fewer attribution errors**. This is a hypothesis to validate, not a claim that competitors lack these capabilities. Avoid making unsupported promises such as perfect accuracy or replacing all human scribe duties.

## 2. Speaker attribution is a system, not a checkbox

**Transcription** recognizes words. **Diarization** groups speech by speaker. **Identification** maps a voice to a person. **Role assignment** maps a participant to clinician, patient, caregiver, or interpreter. These outputs need separate fields and separate evaluation.

For example, a correct word sequence can still produce a dangerous record if a caregiver's statement about their own medication is attributed to the patient. An ASR model can also mistake "I do not take it" for "I do take it." General word accuracy does not adequately measure either problem.

Recommended behavior:

1. Start with anonymous encounter-scoped speaker IDs. Confirm roles with the clinician; never assume the first speaker is the doctor.
2. Show live labels as provisional when the engine can revise them. Stabilize the final transcript using the whole encounter where possible.
3. Maintain explicit speaker mappings across chunks and reconnections. Provider "Speaker 0" from a restarted stream is not automatically the previous Speaker 0.
4. Preserve overlap and unknown attribution. A forced confident label can be worse than a clear request for review.
5. Keep acoustic confidence, speaker confidence, and clinical-review flags separate. A model's confidence is not a calibrated probability of clinical correctness.
6. Let a reviewer correct one turn, a time interval, or a whole participant. Preserve changes and provenance.
7. When a calling SDK supplies separate participant tracks, retain them. Channel separation generally gives stronger attribution evidence than inferring everyone from one mixed microphone, though a track can still contain more than one person.
8. Consider optional clinician reference audio only after defining consent, storage, and deletion. Do not build persistent patient voiceprints for the MVP.

Pyannote's Community-1 documentation provides an offline diarization baseline and discusses an exclusive output that simplifies alignment with transcription. Exclusive output should not erase an overlap warning in our product. [Model documentation](https://huggingface.co/pyannote/speaker-diarization-community-1)

## 3. Speech APIs: capabilities and prices

Rates are public list/displayed rates observed on the research date, excluding taxes, free credits, minimum commitments, storage, retries, support, and negotiated healthcare terms. Per-hour conversions use 60 minutes. A cheap public API tier is not proof that our exact medical workflow is covered by a signed BAA.

| Engine / configuration | Published rate and normalized cost | Fit and limitation |
| --- | --- | --- |
| **Deepgram Nova-3, batch** | **$0.0043/min = $0.258/hour** monolingual; batch diarization included | Initial final-transcript candidate. Nova-3 Medical is a documented model option; its launch page quotes the same starting batch price. Confirm selected medical model, feature, and contract rates. [Pricing](https://deepgram.com/pricing), [Medical model](https://deepgram.com/learn/introducing-nova-3-medical-speech-to-text-api) |
| **Deepgram Nova-3, live + diarization** | Promotional **$0.0048 + $0.0020 = $0.0068/min = $0.408/hour**. Regular base displayed: $0.0077/min, making $0.0097/min with speakers | Keyterm prompting adds $0.0013/min per pass. Do not budget as though the promotional rate is permanent. [Pricing](https://deepgram.com/pricing) |
| **AssemblyAI Universal-3.5 Pro, batch + speakers + Medical Mode** | $0.21 + $0.02 + $0.15 = **$0.38/hour ≈ $0.00633/min** | Medical transcription comparator. Optional keyterms cost another $0.05/hour. [Pricing](https://www.assemblyai.com/pricing) |
| **AssemblyAI Universal-3.5 Pro Realtime + speakers + Medical Mode** | $0.45 + $0.12 + $0.15 = **$0.72/hour = $0.012/min** | Live medical/speaker candidate; keyterms included at this tier. Disable primary-speaker isolation for room conversations unless testing establishes it preserves all participants. [Pricing](https://www.assemblyai.com/pricing) |
| **Speechmatics** | Displayed: Batch Melia 1 **$0.129/hour**; Standard batch/live **$0.24/hour**; Enhanced batch **$0.40/hour**, live **$0.43/hour** | Evaluate accents and live speaker separation. Rates/features vary by model. Pricing has a model-training discount control; obtain a no-training healthcare quote. Enterprise deployment options include private infrastructure. [Pricing](https://www.speechmatics.com/pricing) |
| **ElevenLabs Scribe v2** | Batch **$0.22/hour ≈ $0.00367/min**; keyterms +$0.05/hour, entity detection +$0.07/hour | Batch model documents speaker labels and word timestamps. A strong low-cost comparison candidate, subject to Scribe-specific BAA coverage. [Pricing](https://elevenlabs.io/pricing/api?price.section=speech_to_text), [Capabilities](https://elevenlabs.io/docs/overview/capabilities/speech-to-text) |
| **ElevenLabs Scribe v2 Realtime** | **$0.39/hour = $0.0065/min** | The documented realtime feature list does not establish parity with batch diarization. Do not promise live speaker attribution from the batch feature list. Same contract gate. [Capabilities](https://elevenlabs.io/docs/overview/capabilities/speech-to-text) |
| **OpenAI gpt-transcribe** | **$0.0045/min = $0.27/hour** | Current recommended file model, with keyword/context hints; not a complete diarization solution. [File guide](https://developers.openai.com/api/docs/guides/speech-to-text), [Pricing](https://developers.openai.com/api/docs/pricing) |
| **OpenAI gpt-live-transcribe** | **$0.017/min = $1.02/hour** | Current live model explicitly lacks speaker labels, word timestamps, and confidence scores. Would require a separate attribution/alignment solution. [Realtime guide](https://developers.openai.com/api/docs/guides/realtime-transcription), [Pricing](https://developers.openai.com/api/docs/pricing) |
| **OpenAI gpt-4o-transcribe-diarize** | Estimated **$0.006/min = $0.36/hour** | Existing file diarization endpoint, with optional known-speaker references. Deprecated; not a durable default for a new product. [File guide](https://developers.openai.com/api/docs/guides/speech-to-text) |
| **Amazon Transcribe Medical** | Published examples imply **$0.075/min = $4.50/hour** | Medical dictation/conversation benchmark. Region, language, and model restrictions need checking for the actual deployment. [Pricing examples](https://aws.amazon.com/transcribe/pricing/) |
| **AWS HealthScribe** | **$0.001667/second ≈ $0.10/min = $6/hour**, 15-second minimum | Bundles transcription, roles, entities, and evidence-linked clinical summaries. Broader and costlier than our current scope; documented batch and streaming, U.S. English, N. Virginia. [Pricing](https://aws.amazon.com/healthscribe/pricing/), [Technical scope](https://docs.aws.amazon.com/transcribe/latest/dg/health-scribe.html) |
| **Google Cloud Speech-to-Text V2 / Chirp 3** | Standard public starting rate **$0.016/min = $0.96/hour** | Multilingual/diarization comparator; verify streaming feature availability for the selected language/region. Channels can be billed separately. Legacy medical models are distinct from Chirp 3. [Overview and rate](https://cloud.google.com/speech-to-text), [Chirp 3](https://docs.cloud.google.com/speech-to-text/v2/docs/chirp-model), [Billing](https://cloud.google.com/speech-to-text/pricing) |
| **Azure Speech** | **Unresolved here:** the regional pricing page returned placeholders, not dollar values | Viable streaming/batch and custom speech alternative, especially in Microsoft environments. The page lists realtime diarization as an add-on and batch diarization as included. Obtain a region-specific quote; do not confuse Azure Speech with the Dragon Copilot end-user product. [Pricing](https://azure.microsoft.com/en-us/pricing/details/speech/) |

### Two purchasing details that materially affect the shortlist

**OpenAI migration:** On August 26, 2026, OpenAI announced that `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe`, and `gpt-4o-transcribe-diarize` will be removed on **February 26, 2027**. The replacement transcription models do not establish speaker-feature parity. Keep an independent diarization route if evaluating OpenAI. This hosted API deprecation is separate from the availability of open Whisper weights. [Deprecation notice](https://developers.openai.com/api/docs/deprecations)

**AssemblyAI BAA access:** Its current Medical Mode documentation says upgraded paid accounts can sign a standard BAA from Data Controls at no additional cost. Older support/legal pages direct customers to sales. Verify availability and exact endpoint coverage in the proposed account. This could make AssemblyAI commercially easier to pilot than a vendor requiring an enterprise minimum. Medical Mode supports English, Spanish, German, and French; unsupported languages fall back with a warning. [Medical Mode documentation](https://www.assemblyai.com/docs/pre-recorded-audio/medical-mode), [BAA page](https://www.assemblyai.com/legal/business-associate-agreement)

Deepgram says to contact it for a BAA. ElevenLabs' cited HIPAA documentation concerns **Agents**, enterprise subscriptions, and zero-retention conditions; do not infer it automatically covers every standalone Scribe endpoint. OpenAI likewise requires checking the agreement and endpoint data controls, not treating an API key as healthcare authorization. [Deepgram](https://developers.deepgram.com/trust-security/data-privacy-compliance), [ElevenLabs](https://elevenlabs.io/docs/eleven-agents/legal/hipaa), [OpenAI](https://developers.openai.com/api/docs/guides/your-data)

## 4. Local model options

### Local can mean three different things

| Location | Benefit | Tradeoff |
| --- | --- | --- |
| **On the iPhone** | Works without uploading encounter audio; can work offline after models are downloaded | Battery, heat, memory, model storage, background execution, and device support must be measured |
| **A clinic-owned Mac or server** | Larger models with local operational control; phone can stream over the clinic network | Pairing, network availability, updates, backups, and support become our responsibility |
| **Our private server/cloud GPU** | Shared capacity and controlled model versions; portability | Still a hosted PHI processor; infrastructure, security, and idle capacity are ongoing costs |

### Practical local candidates

| Component | Role | Commercial cost and caveat |
| --- | --- | --- |
| **Apple SpeechAnalyzer / SpeechTranscriber** | On-device streaming and file ASR introduced in iOS 26; designed for longer conversations | No metered ASR service bill for on-device use. Models must be downloaded; check device/language availability. Its reviewed public interface does not supply our complete diarization workflow. [Apple WWDC](https://developer.apple.com/videos/play/wwdc2025/277/) |
| **FluidAudio** | Swift/Core ML ASR, voice activity detection, speaker embeddings, and streaming/offline diarization; Parakeet-family ASR | Apache-2.0 SDK; individual model terms also apply. Especially relevant to the iPhone prototype. Desktop speed claims are not iPhone battery or clinical-quality evidence. [Repository](https://github.com/FluidInference/FluidAudio) |
| **Argmax WhisperKit + SpeakerKit** | On-device Whisper transcription and Pyannote diarization for Apple platforms | MIT open-source SDK plus third-party terms. Pro SDK advertises additional live features and Android support; commercial pricing needs inquiry. [Repository](https://github.com/argmaxinc/argmax-oss-swift) |
| **faster-whisper + WhisperX + Pyannote** | Self-hosted transcription, alignment, and speaker clustering | Open components have no hosted per-minute tariff; budget hardware and operations. Review each code/model license. Better initial fit for a server or clinic Mac evaluation than direct iPhone embedding. [faster-whisper](https://github.com/SYSTRAN/faster-whisper), [WhisperX](https://github.com/m-bain/whisperX), [Pyannote](https://huggingface.co/pyannote/speaker-diarization-community-1) |

For local attribution, FluidAudio documents an offline Community-1 pipeline and streaming alternatives including LS-EEND and Sortformer. Select based on measured speaker stability and overlap behavior; a vendor's general benchmark ranking is not sufficient. [Diarization implementation](https://github.com/FluidInference/FluidAudio#speaker-diarization)

### VoiceStudio disposition

Excluded from the product direction. The initial review established a desktop voice-production workspace, without evidence that it is the best foundation for this medical iPhone workflow. No installation, evaluation, licensing work, or product dependency is planned. [Repository reviewed during initial research](https://github.com/debpalash/VoiceStudio)

## 5. iPhone capture: the constraint to resolve first

| Scenario | Proposed approach |
| --- | --- |
| Doctor and patient in the same room | Record the iPhone microphone in VisitQuill; test positioning and quiet speakers |
| Doctor's post-visit dictation | Use a dedicated single-speaker session or labeled encounter addendum |
| Zoom/Teams running on another device | The phone can hear room speakers, but headphones remove that acoustic path. Prefer a supported platform integration when remote visits become scope |
| Phone/FaceTime/Zoom/Teams call on the same iPhone | Do not promise a generic cross-app audio tap. Use an explicitly supported meeting/telehealth integration, a call routed through infrastructure we control, or user-imported recordings where appropriate |
| Clinician walks through an eight-hour day | Keep the app ready, but use explicit patient/encounter boundaries and pauses. One continuous recording risks mixing patients and increases review and processing costs |

Apple documents that background audio recording can continue with the appropriate audio background mode, including when the screen locks, but calls and other nonmixable sessions can interrupt it. Our engineering conclusion is that uninterrupted microphone capture alongside an unrelated phone call is not a sound product assumption. This requires physical-device testing, not just a simulator demo. [AVAudioSession recording behavior](https://developer.apple.com/documentation/avfaudio/avaudiosession/category-swift.struct/record)

If the remote human scribe is simply listening to an in-person visit through a phone left connected all day, VisitQuill can replace that capture arrangement by recording locally in the app. If the physician is actually consulting patients through that same phone call, the telephony path becomes a first-phase dependency. Confirm this distinction with the customer.

Apple also requires clear recording indication and user consent. A microphone permission prompt is not itself the full patient consent workflow. [App Review Guidelines, recording/privacy](https://developer.apple.com/app-store/review/guidelines/)

## 6. Architecture and future platforms

```mermaid
flowchart TD
  A[iPhone: explicit encounter and consent] --> B[AVFoundation recording and encrypted chunk journal]
  B --> C[Optional on-device live ASR and diarization]
  B --> D[Resumable upload to authenticated service]
  D --> E[Selected contracted ASR and diarization provider]
  C --> F[Versioned transcript and participant mapping]
  E --> F
  F --> G[Speaker correction, uncertainty review, audio playback]
  G --> H[Clinician-approved transcript and controlled export]
  H --> I[Later: clinical note generation and EHR adapter]
```

This is a proposal. In local-only mode, the cloud branch is disabled and cannot activate as a silent fallback.

| Layer | Initial choice | Reason |
| --- | --- | --- |
| iOS UI and audio | SwiftUI, AVAudioSession/AVAudioEngine, Swift concurrency | Direct control of the highest-risk part: recording, interruptions, routes, and native inference |
| Local persistence | Encrypted audio chunks, protected metadata, Keychain-managed secrets | Survive connectivity loss and process termination without losing encounter state |
| Service API | Authenticated REST for records, WebSocket where needed for live updates | Shared contract for later Android, Mac, Windows, and web clients |
| Processing | Provider adapters and queued finalization jobs; Python workers are a practical initial option | Switch engines without rewriting the app or transcript model |
| Records and audio | Tenant-isolated relational database; encrypted object storage with lifecycle rules | Separate access, retention, and revision history for audio and text |
| Operations | Access logs, billing meters, model/version metadata, PHI-free operational telemetry | Investigate failures and understand actual per-practice costs |

SwiftUI is an iOS-first recommendation, not a promise of shared Android UI. Flutter or React Native could share more interface code, but background audio and local inference still require native work and device testing. Keep most reusable value in the service contract, transcript schema, review rules, and evaluation corpus. Android can later use Kotlin/native audio or a shared UI framework. Mac can reuse some Swift modules; Windows can use the same service and a separate capture adapter.

Preserve encounter IDs, provisional patient references, participant roles, timed segments, original model output, human edits, language, consent, and review state. A final EHR integration must independently confirm the patient, destination encounter, and clinician-approved content. We are deliberately building the data boundary now, not implementing chart writes.

## 7. What it costs to serve a doctor

**Illustrative workload:** 20 visits/day × 15 minutes × 22 days/month = **6,600 minutes = 110 audio hours/month**. This is a planning assumption, not measured customer usage. Each full cloud pass processes the entire workload; running live and final models is two passes.

| Configuration | Speech processing / doctor / month | 15-minute visit |
| --- | ---: | ---: |
| Deepgram batch | $28.38 | $0.0645 |
| Deepgram live with speakers, promotional | $44.88 | $0.1020 |
| Deepgram live + complete batch final pass, promotional | $73.26 | $0.1665 |
| Same two passes at displayed regular live rate | $92.40 | $0.2100 |
| AssemblyAI medical batch with speakers | $41.80 | $0.0950 |
| AssemblyAI medical live with speakers | $79.20 | $0.1800 |
| AssemblyAI medical live + complete medical batch pass | $121.00 | $0.2750 |
| ElevenLabs Scribe v2 batch | $24.20 | $0.0550 |
| Speechmatics Enhanced live, displayed base rate | $47.30 | $0.1075 |
| OpenAI gpt-transcribe, without separate diarization | $29.70 | $0.0675 |
| OpenAI gpt-live-transcribe, without separate diarization | $112.20 | $0.2550 |
| Google V2 standard starting rate | $105.60 | $0.2400 |
| Amazon Transcribe Medical | $495.00 | $1.1250 |
| AWS HealthScribe including its additional outputs | Approximately $660.00 | Approximately $1.50 |

Calculations use the cited rates in section 3; they are not vendor quotes. Optional vocabulary features, separate channels, medical contract minimums, and repeated retries can change these amounts. These configurations do not all deliver equivalent outputs or accuracy.

An eight-hour continuous daily stream for 22 days is **176 hours**, or 1.6× this scenario, before extra passes. Do not assume voice-activity detection eliminates all billable silence: metering rules depend on the service and connection behavior.

Additional cost categories:

- Authentication, database, queue, compute, storage, backups, network, and monitoring.
- Clinical quality review, support, onboarding, security work, insurance, and contracts.
- Vendor minimum commitments and reserved concurrency.
- Optional note-generation LLM usage later; no LLM bill is necessary merely to store and display a faithful transcript.
- App distribution and payment fees, customer acquisition, and engineering. These are not part of the speech table.

For a small pilot, reserve an **illustrative $200–$600/month total shared infrastructure budget**, excluding speech APIs, labor, vendor minimums, and legal/security work. This is our planning allowance, not measured infrastructure usage or a cloud quote. Allocate it across actual active clinicians and replace the allowance with a configured bill of materials before pricing a contract.

Local inference avoids a hosted per-minute recognition bill but still has model delivery, device support, QA, and engineering costs. For self-hosting, use:

`effective cost/audio hour = (hardware or server cost + operations) / successfully processed audio hours`

For illustration only, a server assumed to cost $1/hour and kept running 720 hours costs $720/month before operations. Compared with $0.258/audio-hour cloud batch recognition, hardware-only break-even is about 2,791 processed audio hours/month. Whether that server can meet peak concurrency and quality is unmeasured. Cheap cloud ASR can be hard to beat financially at small scale; local processing is initially strongest as a privacy/offline benefit.

### Service pricing hypothesis

Test willingness to pay around **$149–$249/clinician/month**, with explicit included hours and usage-based team plans. This is a proposed discovery range, not a recommendation to launch at those prices. Freed and Heidi already offer clinical documentation at competitive prices; transcription alone needs compelling speaker-review quality or local privacy to justify a premium.

Do not promise unlimited all-day processing before measuring actual costs. At $199 revenue, $73.26 of speech processing and an assumed $40 of other service costs leaves roughly **43% contribution margin**, before engineering and sales. A local live preview plus one cloud finalization pass could improve this, if it passes quality and device tests.

For distribution, direct organizational sales and individual subscriptions can have different App Store purchasing requirements. Apple's enterprise-services exception is specific to sales to organizations/groups; merely describing a subscription as a service does not create a universal exemption. Choose the sales channel before implementing billing. [App Review Guidelines §3.1.3](https://developer.apple.com/app-store/review/guidelines/)

## 8. Healthcare requirements that affect implementation now

U.S. medical audio can contain PHI before any EHR integration exists. HHS states that a cloud provider processing or storing ePHI can be a business associate even if it holds only encrypted data without the key. Signed agreements and appropriate technical/organizational controls belong in the transcription phase. Local inference reduces disclosure but does not alone establish compliance. [HHS cloud guidance](https://www.hhs.gov/hipaa/for-professionals/special-topics/health-information-technology/cloud-computing/index.html)

Design the pilot around clinic-approved recording consent, tenant isolation, encryption, role-based access, clinician authentication, configurable retention, deletion of replicas and vendor copies, auditability, and an incident process. Keep PHI out of analytics, crash payloads, notifications, filenames, and routine support tooling. Review device backup and clipboard behavior. Retain raw audio only for the clinic-approved verification period; do not invent one retention duration for every practice.

The recorded transcript remains a draft until reviewed. Preserve original recognized text when presenting normalized numbers or suggested terminology corrections. Do not let a general-purpose LLM silently rewrite medications, doses, negations, or speaker identities. Keep diagnostic advice, autonomous coding, orders, and chart writes outside this first scope.

## 9. How we establish that it is better

Run a paired evaluation on the same audio, not separate demos with different microphones or speakers. Begin with role-play audio; move to consented clinical audio only within the agreed processing environment. The current human scribe is a useful comparator but is not automatically the error-free reference.

Use a held-out set covering the initial specialty, accented English, quiet/distant speech, masks, room noise, caregiver/interpreter participation, overlapping speech, negation, dosages, units, medication changes, and clinician addenda. Have qualified reviewers independently annotate the reference and resolve disagreements. Keep tuning examples separate from the final test set.

Measure:

- **WER:** overall substitutions, deletions, and insertions relative to reference words.
- **Clinical errors:** medication, dosage, unit, allergy, negation, laterality, and temporal-context errors, adjudicated by severity.
- **Speaker errors:** diarization error rate with stated overlap/collar rules, speaker-attributed word errors, role errors, and label changes across reconnects.
- **Workflow:** median and slowest-decile correction time, replay time, unresolved flags, and clinician preference.
- **Reliability:** recoverable versus lost audio, lock-screen operation, interruptions, Bluetooth changes, app termination, storage exhaustion, and Wi-Fi/cellular handoff.
- **Device behavior:** latency, memory, heat, and battery across the oldest supported iPhone and current devices.
- **Economics:** invoiced minutes, retries, concurrency, full cost per reviewed encounter, and support time.

Initial study proposal: 50–100 varied encounters for discovery, with a separately held-out evaluation portion. A small pilot can identify failure modes; it cannot establish universal clinical safety. Set specialty-specific acceptance thresholds with the clinical lead before evaluating vendors. Treat an unflagged critical speaker/medication error as a release-blocking finding to investigate, not something to hide inside average WER.

Suggested product performance targets, **unmeasured and negotiable**: visible live text within two seconds under defined network/device conditions; final transcript within one minute of stopping a 15-minute visit; every captured audio gap explicitly reported; no silent loss in the agreed recovery tests; materially less clinician correction time than the existing workflow. Quality gates take priority over these latency targets.

## 10. First build sequence

1. **Customer workflow and technical spike:** confirm capture scenario, specialty, devices, language mix, and current scribe duties/cost. Build a native recorder with interruption/recovery instrumentation and run cloud/local comparisons on role-play recordings.
2. **Transcript MVP:** encounter boundaries, participant mapping, recorded audio, live/final states, speaker editor, replay, addenda, and controlled export. Preserve provider-neutral records.
3. **Clinic pilot readiness:** executed data agreements, approved consent/retention, access controls, deletion and recovery verification, support process, and clinician-led quality gates.
4. **Paired customer pilot:** compare clinical and attribution errors, correction effort, and costs; select the production provider/mode and pricing from evidence.
5. **EHR phase:** confirm patient identity and integrate reviewed content. Clinical note generation can be a distinct, evaluated stage.

Planning estimate: a focused technical spike may take 1–2 weeks, followed by several weeks of MVP work and pilot hardening for a small experienced team. Contracting, clinical evaluation, and device failures can dominate the calendar; this is not a delivery commitment.

See the accompanying [transcription MVP blueprint](./transcription-mvp-blueprint.md) for the concrete first-build acceptance criteria.