Maharshi Nahar backend & applied AI Resume

An app that keeps nothing

Kieru gives two people chat, a whiteboard, file transfer, screen share and voice on one screen. When either of them leaves, all of it is gone.

Every chat app is an archive. You send a file to the laptop three feet away and it goes through a data centre in another country, then stays on someone's disk forever. I wanted the opposite: the digital version of talking in a room, where when the conversation is over it is actually over. Kieru is Japanese for to vanish.

"We do not store your data" is easy to say and hard to prove. The rule I gave myself was structural instead: the server may know accounts, friendships, who is online, and each online user's temporary peer ID. Nothing else. If a feature needed the server to see user content, the feature did not ship. Chat, strokes, files and audio go browser to browser over WebRTC and never touch my backend at all.

The hosting is asleep

It runs on Hostinger shared hosting. No Docker, no root, no daemons, and the important one: the Node process sleeps when nobody is using it. Any in-memory state disappears without warning, so there is no module-level cache, no session map, no in-process queue. Every piece of truth lives in MySQL and every request has to be correct on a process that woke up a millisecond ago knowing nothing.

Most of the architecture fell out of that. A signaling server is a long-lived WebSocket server and a TURN relay is a long-lived daemon, so neither can run here; signaling goes to the PeerJS public cloud and relaying to Open Relay's free tier. Presence cannot be a live connection list either, so it is a heartbeat in a table: the client posts its peer ID every 20 seconds and online means a heartbeat inside 45 seconds. That survives cold starts perfectly and costs up to 45 seconds of lag, which is why a closed tab can still look online and why dialing one gives a clean "friend unavailable" instead of pretending it cannot happen.

A stateless API was the right shape for this app regardless. The cheap hosting just removed the option of getting it wrong.

Localhost tells you nothing about P2P

WebRTC needs the two browsers to find a network path to each other, and behind carrier-grade NAT, which is where Indian mobile networks put you, there often is not one. Roughly one connection in seven never goes direct. TURN was a day-one requirement, not an optimisation, because without it the app simply does not work for a slice of users, and it is the slice you cannot reproduce at a desk.

Two things made that testable. First, a force-relay debug flag: iceTransportPolicy: 'relay' makes the browser refuse direct paths, so everything goes through TURN and I find out in ten seconds whether my TURN config works. Without it you cannot tell a correct setup from a lucky direct connection. Second, a rule: no peer-to-peer feature is done until it has run between two different mobile carriers. Two tabs on one laptop prove nothing.

Early on a failed connection just timed out with a useless "did not respond" toast. The fix was to instrument the state machine rather than guess at the symptom: poll iceConnectionState and iceGatheringState every two seconds during a dial, and on timeout print what it was stuck on.

dial timed out. ice was: checking turns an unreproducible complaint into a diagnosis. With WebRTC the error you get is almost never the error you have.

Two ways to blow up a tab with one file

Naive peer-to-peer file transfer fails on both ends independently, and I hit both. On the receiving side, collecting chunks into an array and joining them at the end works beautifully for a 5 MB PDF and kills the tab somewhere past a gigabyte. StreamSaver.js fixes it by handing you a writable stream to the user's disk, so each 16 KB chunk is written and forgotten and memory stays flat whatever the file size.

On the sending side the problem is backpressure. Slice the file and call send() as fast as JavaScript can go and you fill an internal buffer far quicker than the network drains it, until memory balloons and the connection dies. The data channel already has flow control for this: pause when bufferedAmount goes over 1 MB, set bufferedAmountLowThreshold to 256 KB, wait for bufferedamountlow, resume. The send loop becomes await drain, read a slice, send, repeat, and the transfer runs at exactly the speed the link can take on a LAN and on a weak mobile connection alike, with nothing to tune.

sender slice 16 KB data channel buffer pause above 1 MB resume at 256 KB receiver StreamSaver straight to disk bufferedamountlow, send the next slice
Nothing is held at either end. The channel's own buffer sets the pace.

Around that, files are offered rather than pushed, so the receiver sees a name and a size and accepts before a byte moves, and each transfer opens its own labelled data connection so a big file cannot starve chat. There is no resume. Resume means persisting partial state, and this app does not persist. Sometimes the right answer to a feature request is that it contradicts the premise.

The relay is free, which means it is finite

Direct connections cost me nothing, because the bytes never touch my infrastructure. Relayed ones go through a shared metered pool, so one person screen sharing for three hours can burn everyone else's quota. Relayed connections get a daily budget, currently 10 MB of transfer and 20 minutes of screen share per user, and direct connections stay unlimited. Working out which one you are on means calling getStats() after the connection is up, finding the selected candidate pair and checking whether either candidate's type is relay. Most users never see a limit mentioned.

The accounting detail I like: the budget is claimed before the file offer goes out, not after the transfer finishes. The sender reserves the byte count, and if there is no room they are told immediately and the receiver is never even prompted. Reserving afterwards would let someone start ten large transfers at once and blow through the day in parallel. The refund path matters just as much, because without it a declined file would permanently eat someone's allowance.

The daily reset needs no cron job, which is good, because nothing scheduled can be trusted on a process that sleeps. The quota table is keyed on user and day, where day is a plain YYYY-MM-DD string in UTC. Midnight is not an event that has to be handled. The key changes and the new row starts at zero.

Two voice bugs I wrote on purpose

Both voice failures came from adding code, not from missing it. First I re-acquired the microphone after the call connected and swapped it in with replaceTrack(), which felt like hygiene and actually stopped the live track and cut audio a second into every call. Then I checked track.muted before dialing and refused to proceed if it was true, which is a reasonable-looking guard except that muted is routinely true for a moment after getUserMedia and clears itself once samples flow. It blocked calls that would have worked.

Both fixes were deletions. There is a comment in voice.js that exists only to stop me re-adding the first one. I have started being suspicious of code I wrote to be safe.

The real problems needed real code. Safari will not play an audio element without a user gesture, so incoming audio silently does not start; that is handled by catching the play() rejection, toasting "tap anywhere to hear audio" and attaching one-shot listeners that retry. And getUserMedia does not exist at all on insecure origins rather than throwing something useful, so the error message says plainly that the mic needs https.

Reconnection, where the polish went

Networks drop. Someone walks out of wifi range onto mobile data mid-sentence, and handling that well is most of the difference between a demo and something usable. Detection is a ping every 10 seconds over the data channel, and three missed replies means dead, whether or not the browser has fired a close event, because often it has not.

Recovery shows a "reconnecting" overlay without tearing down the session UI, so the user feels a pause and not a crash, then re-fetches the friend's peer ID before retrying with backoff at 1, 2, 4, 8 and 15 seconds. The re-fetch is the step that matters more than it looks: peer IDs are random per login, so if their tab reloaded, the ID I was connected to no longer exists and retrying it would fail forever. There is an abort flag threaded through the whole loop too, because a retry that succeeds after the user hit leave is its own kind of bug.

When it works, the session just continues, and there is nothing to resync because there is no history to resync. Live-only makes reconnection simple in a way an app with message history never gets to be.

What I did not build

No message history, which is the entire premise. No group sessions: two people is one connection, three is three, four is six, and shortly after that you are building a media server, which is exactly the thing this architecture exists to avoid. Not later, no. No video calls, no push notifications, no offline delivery, no file resume.

Almost every rejected feature needed either storing user content or supporting more than two people, which are the two things the app is defined by not doing. That is what a good constraint buys you: most of the scope questions answer themselves.

What I took from it

The first thing I built was not login or the UI. It was two phones on different mobile networks sending each other the word hello over WebRTC through TURN, because if that had not worked, nothing after it mattered. Ordering work by risk instead of by dependency is the habit I would keep from this project.

The other one is smaller and I keep relearning it. The platform already had the flow control I needed, the browser already exposed the state I was guessing at, and both voice bugs were my own defensive code. Roughly 3,200 lines of frontend and backend together, and the parts that took longest were the ones where I stopped writing and started reading what was already there.