Building a Browser Voice Agent: Consent, Workers, and Tool Boundaries
A code-backed look at WebML’s voice agent: what loads on demand, where work runs, and where the current tool boundary still needs tightening.
WebML's voice agent is a useful stress test for browser AI because it combines a microphone, multiple model downloads, a planning model, and tools that can change the page. The interesting engineering work is in the seams: when to request permission, how to recover when a model fails, and which actions a generated plan should be allowed to take. This is a walkthrough of the implementation in the source tree, with its limitations left visible.
1. Keep the first page load light
The voice interface starts as a small trigger. The full provider and its workers are loaded only after the visitor chooses to start. The consent card shows an approximate download size before any large model files are fetched. This matters more than a fast animation: the models need substantial bandwidth and storage, and a reader who came for an article should not pay that cost.
The browser separately asks for microphone permission when capture begins. Dismissing the card does not start recording. If you build a similar widget, check both boundaries: model download consent and the browser's microphone permission are different decisions.
2. Separate capture, inference, and playback
An AudioWorklet copies microphone samples and transfers each buffer to the speech worker. The provider owns the media stream and stops its tracks during cleanup. Separate workers hold the speech recognizer, planner, and speech synthesis model. The UI receives progress and output messages while the heavier work runs away from the rendering thread.
Workers protect the page from synchronous model work, but they do not make GPU memory unlimited. Two large models can still compete for the same device resources. A practical failure path is therefore part of the design: report a load error, let the visitor reset, and leave the rest of the site usable without the voice feature.
3. Treat an agent plan as untrusted input
The planner can request site actions such as search, read, navigate, and scroll. Each maps to a named WebMCP tool with an input schema. The adapter looks up the named tool at execution time and returns an error if it is absent or rejects its arguments. Keeping that boundary explicit prevents a malformed model response from being treated as a valid site command.
There is an important exception in this repository: the voice agent also exposes a JavaScript action for page interaction. Generated JavaScript has much broader reach than a narrow read or navigation tool. The site's WebMCP local_jstool is restricted to development, but the agent's separate JavaScript action needs the same scrutiny. Do not reuse this pattern for a privileged workflow without removing that action or adding a clear approval boundary. The safest next step here is to narrow the production tool set.
4. Tell visitors where the network is still involved
The computation runs locally, but the first visit downloads models. The site also loads analytics. Local inference describes where prompts are processed; it does not mean the whole site is offline or that no network requests occur. Model caches are controlled by the browser and can be evicted. A cached model may speed a return visit, but it is not an offline guarantee.
For a smaller experiment with a simpler path from model selection to output, try the LLM Playground. For the exact disclosures that apply across this site, read the privacy page.