See the other ways to buy On-site & off-net

Voice agents that never leave your building

One or two Dell rack-mount servers, delivered to your data center, running the same agent back end as our cloud platform. Speech recognition, speech synthesis, and the language model all run on GPUs in your rack — no Internet calls, no audio or patient data leaving your network.

Built for HIPAA-sensitive and air-gapped environments

A short form and one of our engineers calls you back — every deployment includes 100 hours with our integration team.

When the calls can’t go off-net

Some conversations can never touch a third-party API. Clinics, labs, public safety, finance, defense suppliers — if that is you, the agent comes to your rack instead.

No Internet calls

Every model runs on the server in the rack. Nothing is sent to an outside provider for processing.

HIPAA-friendly by design

Audio, transcripts, and PHI stay inside your security perimeter, under your existing controls and audits.

Latency you control

The agent sits on the same LAN as your PBX. No public network between a caller’s question and the answer.

Keeps working off-net

An upstream outage doesn’t take your agent down. Internal calls keep being answered.

What arrives on the dock

One rack unit, or two

We specify, build, and burn in the hardware before it ships, then rack and commission it with your team.

Dell rack-mount servers

A single Dell server handles most sites. Choose a dual-server configuration when you want redundancy for the agent or more concurrent calls than one chassis can carry.

  • Enterprise hardware with the Dell support contract you already know.
  • Standard rack, power, and network requirements documented up front.
  • Dual configuration adds headroom and a second node for failover.

GPUs sized to your call load

Speech and language models want GPU. We offer a range of accelerator cards and pick from it based on the concurrent call volume you expect — not on a one-size-fits-all bundle.

  • Entry cards for a handful of simultaneous calls, larger cards for busy call centers.
  • We size the card to your busiest hour, then leave room to grow.
  • Add or upgrade cards later as call volume and model choices change.

The management UI, served locally

The same web interface your team would use on our cloud platform runs on the appliance. Your administrators open it from the LAN to manage agents, prompts, and call records.

  • Manage agents, context, and integrations from your browser on-net.
  • Call records and transcripts stay on the appliance with the rest of your data.
  • SIP-native, so it registers to your PBX like any other extension.
The same stack, in your rack

Not a stripped-down edition

The agent back end on the appliance is the same one that runs VoxLayer. The same speech recognition, text-to-speech, and language model choices you would pick in our cloud are available locally — and every one of them executes on the server in your rack.

Speech in and out, locally

Your ASR and TTS selections run on the appliance GPUs. Caller audio is transcribed and the reply is spoken without a packet leaving the building.

Local language models

Pick from the same model line-up and host it on the appliance. No external LLM API, no per-token bill, no prompts in someone else’s logs.

Off-net by default

The appliance needs no outbound Internet access to answer a call. If your security team wants it isolated, it can be.

Bring your own agent

Or use it purely as a SIP, ASR, and TTS gateway

Already have an agent, or need one whose every word you control? Let the appliance own the telephone call and the speech, and put your own program in the middle.

We manage everything on the telephony side — SIP registration to your PBX, call setup and teardown, transfers, barge-in — plus speech recognition and speech synthesis on the appliance GPUs. Then, instead of calling our language model, the appliance opens a WebSocket to your agent, sends it each caller utterance as a transcription, and speaks whatever text you send back.

  • We keep the hard parts — the PBX integration, the real-time audio pipeline, interruption handling, and the ASR and TTS models your callers hear.
  • You keep the conversation — any language, any framework, on your own server. Use your own model, a rules engine, or no model at all.
  • Total control of what it says and sees — nothing reaches a caller that your code did not write, and the agent touches only the systems you hand it.
  • A protocol you can read in a minute — a CALL frame with the caller details when the phone is answered, a USER line for every utterance, and plain text back for the agent to speak.
your agent, over the wire
appliance your agent
 CALL {"caller_extension":"1042","dialed_number":"5551234","agent_id":61,"channel":"voice"}
 USER pawn to e4
 Bold. I'll answer knight to f6. Your move.
 USER bishop takes f7
 That's check, but you left your queen hanging.

Start from the chess demo

We ship a hundred-line PHP file that plays a game of chess over a phone call. It is the whole protocol, end to end, in something you can read in one sitting — and a working starting point for your own agent.

Included with every deployment

100 hours with our integration engineers

Hardware is the easy part. The deployment includes a hundred engineering hours to make the agent genuinely useful on your phones — and to hand it over to your team.

1

Design the call flow

We sit down with the people who answer your phones today. What the agent handles, what it asks, when it transfers, how it escalates after hours — mapped out and built, not guessed at.

2

Integrate your local applications

The agent is only as useful as what it can reach. We connect it to the systems already running on your network — scheduling, records, ticketing, directory, whatever the call needs — over your internal interfaces.

3

Train your staff engineers

Your engineers learn to run it: day-to-day maintenance, monitoring, model and prompt changes, and first-line support — so the appliance is yours to operate, not a black box.

After go-live

We stay your AI partner

Commissioning day is the start, not the finish. As you find new things the agent should do, we are still here to build them — new capabilities, new integrations, new agents for other departments.

Your team handles the day-to-day after training. We handle the harder additions and keep the platform on the appliance current with what we ship in the cloud.

New agent skills

Added as your use cases grow beyond the original call flow.

Further integrations

More of your internal systems wired in as you need them.

Ongoing knowledge transfer

New staff engineers get up to speed on the platform.

Engineer-level support

Escalation to the people who build the platform, not a script.

Live demo

Questions? Press the button and ask

Our demo agent has been briefed on this product specifically. If anything on this page raises a question — GPU sizing, what the 100 hours cover, how it stays off-net — start a conversation and ask it directly.

  • Same voice pipeline as a live phone call.
  • Low-latency, natural back-and-forth conversation.
  • Runs in your browser — nothing to install.

Live demo is not wired yet. On agent 61, add https://voxlayer.ai/on-site under Inbound → Web (use existing button, selector .talk-btn-onsite), then save.

Live demo coming soon

Our live in-browser agent is being prepared. Create a free account to launch your own agent and test it instantly.

Try it with a free account

Let’s scope your deployment

Tell us how to reach you and an engineer will call back to talk through call volume, server and GPU sizing, the applications you want integrated, and what your compliance team will need to see.

Four fields. No obligation, no drip campaign.