One or two Dell rack-mount servers, delivered to your data center, running the same agent back end as our cloud platform. Speech recognition, speech synthesis, and the language model all run on GPUs in your rack — no Internet calls, no audio or patient data leaving your network.
A short form and one of our engineers calls you back — every deployment includes 100 hours with our integration team.
Some conversations can never touch a third-party API. Clinics, labs, public safety, finance, defense suppliers — if that is you, the agent comes to your rack instead.
Every model runs on the server in the rack. Nothing is sent to an outside provider for processing.
Audio, transcripts, and PHI stay inside your security perimeter, under your existing controls and audits.
The agent sits on the same LAN as your PBX. No public network between a caller’s question and the answer.
An upstream outage doesn’t take your agent down. Internal calls keep being answered.
We specify, build, and burn in the hardware before it ships, then rack and commission it with your team.
A single Dell server handles most sites. Choose a dual-server configuration when you want redundancy for the agent or more concurrent calls than one chassis can carry.
Speech and language models want GPU. We offer a range of accelerator cards and pick from it based on the concurrent call volume you expect — not on a one-size-fits-all bundle.
The same web interface your team would use on our cloud platform runs on the appliance. Your administrators open it from the LAN to manage agents, prompts, and call records.
The agent back end on the appliance is the same one that runs VoxLayer. The same speech recognition, text-to-speech, and language model choices you would pick in our cloud are available locally — and every one of them executes on the server in your rack.
Your ASR and TTS selections run on the appliance GPUs. Caller audio is transcribed and the reply is spoken without a packet leaving the building.
Pick from the same model line-up and host it on the appliance. No external LLM API, no per-token bill, no prompts in someone else’s logs.
The appliance needs no outbound Internet access to answer a call. If your security team wants it isolated, it can be.
Already have an agent, or need one whose every word you control? Let the appliance own the telephone call and the speech, and put your own program in the middle.
We manage everything on the telephony side — SIP registration to your PBX, call setup and teardown, transfers, barge-in — plus speech recognition and speech synthesis on the appliance GPUs. Then, instead of calling our language model, the appliance opens a WebSocket to your agent, sends it each caller utterance as a transcription, and speaks whatever text you send back.
→ CALL {"caller_extension":"1042","dialed_number":"5551234","agent_id":61,"channel":"voice"} → USER pawn to e4 ← Bold. I'll answer knight to f6. Your move. → USER bishop takes f7 ← That's check, but you left your queen hanging.
We ship a hundred-line PHP file that plays a game of chess over a phone call. It is the whole protocol, end to end, in something you can read in one sitting — and a working starting point for your own agent.
Hardware is the easy part. The deployment includes a hundred engineering hours to make the agent genuinely useful on your phones — and to hand it over to your team.
We sit down with the people who answer your phones today. What the agent handles, what it asks, when it transfers, how it escalates after hours — mapped out and built, not guessed at.
The agent is only as useful as what it can reach. We connect it to the systems already running on your network — scheduling, records, ticketing, directory, whatever the call needs — over your internal interfaces.
Your engineers learn to run it: day-to-day maintenance, monitoring, model and prompt changes, and first-line support — so the appliance is yours to operate, not a black box.
Commissioning day is the start, not the finish. As you find new things the agent should do, we are still here to build them — new capabilities, new integrations, new agents for other departments.
Your team handles the day-to-day after training. We handle the harder additions and keep the platform on the appliance current with what we ship in the cloud.
Added as your use cases grow beyond the original call flow.
More of your internal systems wired in as you need them.
New staff engineers get up to speed on the platform.
Escalation to the people who build the platform, not a script.
Our demo agent has been briefed on this product specifically. If anything on this page raises a question — GPU sizing, what the 100 hours cover, how it stays off-net — start a conversation and ask it directly.
Live demo is not wired yet. On agent 61, add https://voxlayer.ai/on-site under Inbound → Web (use existing button, selector .talk-btn-onsite), then save.
Our live in-browser agent is being prepared. Create a free account to launch your own agent and test it instantly.
Try it with a free accountTell us how to reach you and an engineer will call back to talk through call volume, server and GPU sizing, the applications you want integrated, and what your compliance team will need to see.
Four fields. No obligation, no drip campaign.