· 13 min read

How to Build a Video Chat App in 2026: Architecture, Vendors, Steps

A video call works on the first try in a demo. It fails in the real world: on hotel Wi-Fi, behind a hospital firewall, on a three-year-old Android phone in a moving car. Whether you are adding video to a telehealth, learning or sales product or building a meeting tool for one industry, the hard part is not getting faces on a screen. It is keeping the call alive and the data handled correctly. Here is how to choose the architecture and build it in the right order.
A laptop with a video call grid, a phone with an incoming call and a media server connecting them, illustrating how to build a video chat app
Adding video to your product?Tell us who joins the calls, on which devices, and what your industry requires. We’ll suggest the build model, a vendor shortlist and a first version.
Scope my video feature

Start with three decisions, not codecs

Vendors will happily talk about codecs and resolutions. Three plainer questions decide what you build:

  1. Who is on the call? Two people (a doctor and a patient), a small group (a tutor and five students) or a large room (a 300-person town hall). Room size picks the topology.
  2. What happens to the media? Calls that are recorded, transcribed or summarized need servers that can see the video. Calls that must never be stored point the other way.
  3. Which rules come with your users? Patients, children, students and recorded sales calls each bring their own laws. More on this below.

With those answers, you can pick one of four build models. Each one moves work between you and a vendor:

  • Examples: telehealth visits, tutoring sessions, support calls with screen sharing, video interviews inside a hiring tool.
  • You rent: media servers, TURN relays, global regions, client SDKs for web, iOS and Android, often recording and captions.
  • You build: users, rooms, permissions, the waiting room, the call screen, failure handling and everything that happens before and after the call.
  • Watch out: per-minute pricing and vendor changes. Keep the SDK behind your own interface so a switch is a project, not a rewrite.

One more decision hides inside the first question: topology. Peer-to-peer sends video directly between devices and suits one-to-one calls. An SFU (selective forwarding unit) receives each stream once and forwards it to everyone else, which is how almost every group call works today. An MCU mixes all streams into one picture on the server: heavy on servers, but useful for recordings and phone dial-in.

How to build a video chat app, step by step

Nine steps in the order that avoids the most rework. Steps 1 to 3 take a few weeks and decide most of the cost.

  1. Write down the call shapeTypical and maximum participants, average length, calls per month, devices and browsers. "Two people, 20 minutes, mostly phones" and "eight people, an hour, laptops on office networks" lead to different products and different bills.
  2. Map the rules for your verticalHealth data, children, student records and recorded calls each bring requirements: signed agreements with vendors, consent screens, retention periods. Confirm them with counsel before you choose a vendor, because they narrow the list.
  3. Choose the build model and vendorCompare two or three SDKs on your call shape: regions near your users, mobile SDK quality, recording and caption options, the contract terms your vertical needs and the price at your expected minutes. Build a one-week spike on real phones before you sign.
  4. Design the backend around rooms and tokensYour server decides who may join which call and issues a short-lived token for that room only. The video vendor never decides access. Rooms, participants and call events live in your database.
  5. Design the call for failureA pre-call device and network check, a waiting room, clear states ("reconnecting", "the other person left"), audio-only fallback and a way back into the call. Plan the screens for a bad connection first, then the good one.
  6. Build mobile with the platform rulesIncoming calls on iOS use CallKit and VoIP push; Android needs a foreground service and its own call notifications. Handle a phone call interrupting a video call, Bluetooth headsets and the app going to the background.
  7. Add recording, captions and AI only where neededEach one creates stored data you must protect, retain and delete on time. Turn them on per room type, with consent, rather than for everyone.
  8. Test on bad networks and real devicesThrottle bandwidth, add packet loss, switch from Wi-Fi to cellular mid-call, join from behind a corporate firewall. Automate what you can; keep a set of older phones for the rest.
  9. Launch small and watch call qualityCollect per-call stats: join time, failed joins, freezes, how often calls fall back to TURN or audio. Fix the worst device and network combinations before you open up.

What you build and what you rent

With an SDK, the media layer is rented. Everything a user remembers about the call is yours. A typical split:

LayerUsually rentedUsually built
Media serversThe SDK vendor’s SFU network across regionsA thin interface over the vendor, so you can switch
ConnectivitySTUN and TURN relays for networks that block direct trafficClear errors and fallbacks when a network still blocks the call
AccessToken signing helpers in the vendor SDKUsers, roles, rooms, who may join, kick, mute or record
Call experienceSample UI components, sometimes a prebuilt call screenWaiting room, device check, layouts, screen sharing, in-call chat
Recording and captionsCloud recording, transcription and caption servicesConsent, storage location, retention, who can watch or download
Around the callCalendar, email and SMS services, paymentsScheduling, reminders, notes, follow-ups, the business workflow

Video minutes are cheap at list prices: a 20-minute call between two people uses 40 participant minutes, about 15 cents. The contract tier your vertical needs, recording storage and TURN traffic usually cost more than the minutes. Our telemedicine app development cost guide compares Twilio, Vonage and Zoom pricing for health use cases.

Architecture for calls that survive bad networks

Most video products that get rebuilt made the same early choice: they trusted the happy path. These rules come from what breaks first in production.

The core of a video chat product (simplified) Client A web or mobile app Client B strict firewall Your backend users, rooms, tokens TURN relay when direct fails SFU media server forwards each stream Recording, captions, AI Storage retention rules Solid: media. Dashed: sign-in, signaling, access. With an SDK, the green and orange boxes are rented.
Your backend decides who joins which room. The media servers only carry video, and a TURN relay rescues the calls that firewalls would otherwise block.
  • Your server owns access. Issue short-lived tokens per room and per role. Never put vendor API keys in the app, and never let a guessable room name be the only lock.
  • Plan for TURN from day one. Corporate, hospital and school networks often block direct connections. A TURN relay carries the media instead. It works over standard web ports, so calls get through, and you pay for the relayed bandwidth.
  • Adapt quality per viewer. Simulcast or scalable video lets the SFU send a small stream to a weak phone and a sharp one to a laptop. Prioritize audio: people forgive a frozen picture, not broken speech.
  • Budget bandwidth per minute. A 720p stream is roughly 1–2 Mbps. Multiply by participants and minutes and you know what self-hosting will cost in traffic, and what your users need on their side.
  • Choose encryption with eyes open. WebRTC encrypts media in transit by default. End-to-end encryption on top of an SFU hides the media from your own servers, which also disables server-side recording, captions and AI notes.
  • Measure every call. Log join time, packet loss, freezes and TURN usage per call. Without that data, "the call was bad" tickets can’t be fixed.

Scaling follows the topology. One SFU server handles a limited number of streams, so large rooms spread across several servers, and audiences above a few dozen active participants often move to a broadcast model. If your product is one person talking to thousands, read our guide on how to build a live streaming app instead.

Compliance by vertical

The same video call carries different obligations depending on who is on it. The common cases for US products:

VerticalWhat usually appliesWhat it changes in the build
TelehealthHIPAA: a Business Associate Agreement with every vendor that touches health dataA vendor plan that offers a BAA, recording off by default, access logs, encrypted storage
EducationFERPA for student records; COPPA for users under 13Parental consent flows, limits on what is recorded and who can view it, school data agreements
Sales and supportState call-recording laws; some states require consent from everyone on the callClear recording notices and consent, pausing recording when card numbers are read out (PCI DSS)
Meeting tools for businessBuyer security reviews, often SOC 2 reportsSingle sign-on, admin controls, data location options, audit logs
All productsAccessibility expectations (ADA, WCAG)Captions, keyboard and screen-reader support for every call control
General information, not legal advice

Rules depend on your users, states and contracts. Confirm your setup with counsel who knows your industry before you choose a vendor or turn on recording.

What goes into the first version

A first version of a video product should be narrow on features and thorough on failure handling. A typical split:

At launchCan wait
One call type that matches your main use caseWebinars, breakout rooms, large events
Pre-call device and network check, waiting roomVirtual backgrounds and filters
Reconnection, audio-only fallback, clear call statesCustom layouts and branding per customer
Screen sharing on desktop, in-call chatWhiteboards, co-browsing, file collaboration
Role-based access and short-lived room tokensPhone dial-in
Call quality logging and an admin view of callsAI summaries and searchable transcripts
Recording only if the workflow needs it, with consentRecording for every call by default

For many products the first version is a web app, because a link that opens in the browser removes the install step for guests. Add native mobile apps when users need incoming call alerts, background audio or a home screen presence.

Built by Gilzor

Results we’ve shipped

70+products launched
98%delivered on time
85%clients come back
Art Scherbakov, Co-FounderAndrew Laminsky, CTOYuri Rudenya, Head of Mobile Development at GilzorAlena Timofeeva, Product Marketing Lead

Talk to the people who build it. Tell us about your project and get a free estimate of scope, timeline and cost.

See how we’d approach yours

Timeline and budget at a glance

$0.0035–0.0041Per participant minute at 2026 list prices of major video APIs
$20–45kEngineering for a production video module on an SDK, Central European rates
+$60–150kWhat self-hosting can add before it handles real-world networks well
3–5 moA focused telehealth visit app, from discovery to launch

There is no single price for a video chat app, because the video sits inside a product. A simple app costs about $40k–90k and a mid-complexity product with a custom backend and integrations $90k–250k, as our app development cost guide shows. A single-practice telehealth app runs roughly $60k–130k; see the telemedicine app development cost breakdown. For one-to-many video, the live streaming app development cost guide has the ranges and the cost per viewer-hour.

Mistakes that cost the most later

  • Choosing peer-to-peer for group calls. It looks free in testing with three colleagues. At five or six participants on phones, uploads saturate and calls fall apart. Moving to an SFU later means rewriting the call layer.
  • Skipping TURN. Calls work in the office and fail for the hospital, school or enterprise customer behind a strict firewall. Those are often the customers who pay the most.
  • Calling the vendor SDK from every screen. When pricing, terms or the product itself changes, you rewrite the app. Twilio announced the end of its Programmable Video product in 2023 and reversed the decision in 2024; teams with their own interface had options either way.
  • Recording everything "just in case". Every stored call is data to protect, retain and delete on schedule, and in telehealth it is health information. Record by purpose and with consent.
  • Testing only on fast Wi-Fi and new phones. Your users join from cars, cafes and old Android devices. Network throttling and a shelf of older phones are cheaper than a month of bad reviews.
  • No call quality data. Without per-call metrics, support can only say "try again". Logging join time, freezes and fallbacks turns complaints into fixes.

Video readiness checklist

Tick what is already true for your product. It shows how ready your video feature is for real users on real networks.

Video feature readiness

FAQ

How do you build a video chat app?
Start with the call shape: how many people join, on which devices, and whether calls are recorded. Then pick a build model. Most teams use a video API or SDK (Twilio, Vonage, Zoom Video SDK, Agora, LiveKit Cloud and similar) and build their own backend for users, rooms, access tokens and permissions. Add a pre-call device check, reconnection and audio-only fallback, test on real phones and weak networks, and launch to a small group while you watch call quality data. Self-hosting an open-source media server comes later, if volume justifies it.
What is the difference between P2P, SFU and MCU?
In peer-to-peer (P2P) calls, every participant sends video directly to every other participant. It is cheap and private, but it breaks down beyond three or four people because each phone uploads several streams. An SFU (selective forwarding unit) is a server that receives one stream from each participant and forwards it to the others, choosing a quality that fits each viewer. It is the standard for group calls today. An MCU (multipoint control unit) mixes all streams into one video on the server. It is easy on weak devices and useful for recording or phone dial-in, but it costs much more server power and adds delay.
How long does it take to build a video calling app?
A new product with video at its center, such as a focused telehealth visit app, typically takes 3–5 months from discovery to launch. Adding video to an existing product on an SDK is shorter, because the call module is only part of that work. Self-hosting your own media servers adds months of infrastructure and testing on top.
How much does it cost to build a video conferencing app?
It depends on the product the video sits in. A production video module on an SDK costs about $20k–45k of engineering at Central European rates, and self-hosting can add $60k–150k before it handles real-world networks well. The whole product lands where any app does: about $40k–90k for a simple one and $90k–250k for a mid-complexity product with a custom backend and integrations. A single-practice telehealth app runs roughly $60k–130k. Usage is cheap: about $0.0035–0.0041 per participant minute at list prices.
Should I use a video SDK or build on WebRTC myself?
Use an SDK unless you have a clear reason not to. Raw WebRTC is free and works well for one-to-one calls, but you still need signaling, TURN servers, monitoring and fixes for every browser and phone. Open-source SFUs such as mediasoup, Janus, Jitsi or LiveKit remove license fees, not work: you run the servers, scale them across regions and stay on call. That trade starts to pay off at very high minute volumes or when you need deep control over the media, and many teams start on an SDK and migrate later behind their own interface.
Can a video chat app be end-to-end encrypted?
Yes, but it costs features. WebRTC always encrypts media in transit, yet an SFU can technically see the media it forwards. True end-to-end encryption adds a second layer that only participants can decrypt, using browser APIs for encoded frames on the web and SDK support on mobile. Once the server cannot see the media, server-side recording, live captions, AI summaries and phone dial-in stop working or must move to the participants’ devices. Decide which matters more for your users before you build.

Where Gilzor fits

We build web apps and native and cross-platform mobile apps, the backends and integrations behind them, and the QA that tests them on real devices. Only 5% of the tasks our developers send to QA come back to them. Our healthcare and edtech pages show the industries where video calls matter most. We work from Poland and Cyprus, with a few shared hours a day with the US East Coast.

Send us who joins your calls, on which devices, how many minutes you expect a month and which industry rules apply. We'll help you choose between an SDK and self-hosting and cut a first version that holds up on real networks.

No sales pitch

Get a straight answer for your project

Tell us what you’re building. We’ll reply with options, a rough cost and timeline. If we’re not the right fit, we’ll say so.

Next, a few optional questions so the first call is useful. We use your details only to reply to your request. Privacy Policy

Andrew Laminsky
Written byAndrew Laminsky

CTO of Gilzor. Responsible for architecture and the engineering standards our teams work by.

LinkedIn →

Gilzor · Web Development partner

Need a team for your web product?

95%referred by business partners
70+successful launches
85%repeat business
98%delivered on time

The team behind them

Art Scherbakov
Art ScherbakovCo-Founder
Andrew Laminsky
Andrew LaminskyCTOLinkedIn
Yuri Rudenya
Yuri RudenyaHead of Mobile Development at GilzorLinkedIn
Alena Timofeeva
Alena TimofeevaProduct Marketing LeadLinkedIn
Tell us what you’re buildingOptions, a rough cost and timeline for your project. No commitment.

More insights