Skip to content

Build AI that shipsMeetJariwala

Built to run, not to demo

PROOF

Built for production

An inference server built for the hot path

A low-latency service for hosting AI models, strictly typed end to end, with the request path kept deliberately short.

Fast enough to feel instant

Turn detection that calls the end of a turn from the shape of the audio, instead of waiting out the silence.

Audio read at the rate it arrives

Reading the turn boundary out of the stream as it comes in, without falling behind it.

Rust where the latency budget is

Rust and Tokio on the hot path, ONNX and Candle running the model, Python where it speeds the work up.

Every number here is measured

The figures on this page come off systems that are running, not out of a deck. Each one is instrumented where it actually serves traffic.

It ships, or it does not count

Everything here reached production and stayed there. No prototypes kept alive for a portfolio, no demos that only run locally.

Meet Jariwala

More about me

Everything I actually work in.

Speech Domain, Models at scale: WHISPER VOICE AI DIARIZATION VAD END-OF-TURN STT TTS VOICE-CLONING Inference, Serving at scale: CONCURRENCY INFERENCE SERVER LLM THROUGHPUT LATENCY STREAMING REAL-TIME Making it fast, Optimization at scale: CUDA GPU QUANTIZATION ONNX TENSORRT METAL GGUF SAFETENSORS Languages, What I build in: PYTHON RUST Over the wire, Transport and queues: gRPC WEBSOCKETS HTTP STREAMING Retrieval, Context engineering: RAG EMBEDDINGS VECTOR DB RANKING MEMORY COMPRESSION Running in prod, Infrastructure: DOCKER POSTGRES REDIS MONITORING EVALS PIPELINES Shipped work, Products, not demos: VOICE AGENTS INFERENCE SERVER TURN DETECTOR MODELS DOCUMENTPORTAL MILLION HOURS DATASETS Craft, How it feels to use: UI WEB WORKFLOW STORYTELLING SYSTEMS Off the keyboard, The other half: SPEAKER WRITER MENTOR CREATOR BUILDER About, Meet Jariwala: AI THAT SHIPS SYSTEMS PROOF SCALE SHIPPING

Selected work

FAQ

AI systems end to end: real-time inference, audio and ML pipelines. The bar is simple, it has to run in production, not just demo well.

Let’s build something that ships.

Let's talk