Self-hosted speech and voice AI
Voice is a different kind of sensitive data from text. A call recording carries a person's actual voice, which some jurisdictions treat as biometric data, plus whatever they say, unfiltered in a way a written message rarely is, since people say things on a call they wouldn't write down. Send that audio to a third-party speech API for transcription or a voice assistant pipeline, and the raw recording leaves your infrastructure, not a summary of it.
The two workloads: transcription and voice assistants
Transcription is the simpler case: audio in, text out. Speech-to-text has become one of the more mature open-weight AI categories, and the resulting text can then feed into anything text-based, a meeting summarizer, a support ticket, a searchable call archive, including the kind of retrieval setup covered in self-hosted semantic search.
A voice assistant is a pipeline built on top of that: speech-to-text turns what someone says into text, an LLM decides how to respond, and text-to-speech turns the reply back into audio. Each stage is a separate model, and each one can run locally, meaning the entire round trip, someone's actual voice, the words they said, and the reply generated for them, never has to leave the machine it runs on.
What we can say about running this on a Spark
We haven't published benchmarked throughput numbers for speech models specifically, the way we have for text-generation benchmarks, so treat the following as a description of what fits rather than a measured performance claim. A speech-to-text model, an LLM for the assistant's reasoning and reply text, and a text-to-speech model can all coexist within a single Spark's 128GB of unified memory, the same way an embeddings model, vector database, and LLM coexist for a document search setup. Unified memory is what makes that combination practical on one box: three separate model types, each with its own memory footprint, share the same pool rather than needing three separate GPUs sized for the largest one.
What that means in practice depends heavily on your specific setup, model sizes, audio quality, expected latency, and concurrent call volume, all move the answer around, so if throughput or latency at a specific call volume is a hard requirement, that's worth testing against your own workload rather than assuming a number. We'd rather say that plainly than publish a figure we haven't measured.
Why self-hosting matters more for voice than text
A support call, a sales call, an internal meeting recording, all of these usually include people who never explicitly agreed to have their voice sent to a specific AI vendor for processing, even if they were told the call "may be recorded for quality purposes." That's a lower bar than informed consent to a specific third party handling the audio. Running speech-to-text and any downstream LLM processing on infrastructure your organization controls keeps that audio inside the boundary implied by "this call may be recorded," rather than extending it to whichever API vendor happens to sit behind the transcription feature.
The same logic applies to internal voice assistants, an internal tool employees talk to for scheduling, IT requests, or quick lookups. Voice input to that kind of assistant often includes incidental context the employee didn't mean to share as data, a name mentioned in passing, background conversation picked up before they realized the mic was live. Keeping the whole pipeline local means that incidental audio never becomes a third party's log file.
Getting started without overcommitting
Transcription is the easier entry point than a full voice assistant, since it's a single model with a well-understood interface, audio file in, text out, and it slots into an existing workflow rather than requiring you to design a whole conversational pipeline up front. Start there: run meeting recordings or support call audio through a self-hosted transcription model, feed the output into whatever text-based tooling already exists, a summarizer, a searchable archive, and see whether transcript quality and turnaround meet the bar before investing in the full speech-to-text-to-LLM-to-speech loop a voice assistant needs.
Audio quality going in affects output quality more than most people expect. A clean single-speaker recording transcribes far more reliably than a multi-speaker call with crosstalk and background noise, and no amount of downstream LLM cleanup fully compensates for a bad source recording. If accuracy on real call audio matters, test with your own recordings rather than a demo clip before building a workflow around the result.
Where the hardware fits
A single Spark at $0.79/hour in EU-Central is enough to run a speech-to-text model, an instruct model for reasoning, and a text-to-speech model together for most internal use cases, transcribing meetings or calls, or running a voice assistant for a small-to-mid-size team. For call-center-scale concurrent volume, where many simultaneous audio streams need low-latency transcription at once, a multi-node cluster gives the memory and throughput headroom a single box won't have; talk to us about sizing that against your actual call volume rather than guessing from a single-Spark baseline.