AI Engineer & Backend Developer at Purple Sky Infotech
I build and ship production AI systems end-to-end — real-time voice agents, RAG pipelines, and self-hosted LLM infrastructure.
About
I'm a software engineer working at the intersection of AI and backend systems. Right now I'm building a high-throughput, “true natural” voice-agent platform at Purple Sky Infotech — orchestrating STT/TTS engines and LLM APIs to hold real conversations at sub-second latency.
I'm drawn to the hard parts: self-hosting open-source LLMs on GPUs with vLLM, cutting inference cost without sacrificing latency, and building RAG systems that stay fast over hundreds of thousands of document chunks. I care about systems that are correct, observable, and cheap to run.
Education
B.E. in Electronics & Communication Engineering
Jawaharlal Institute of Technology, Vidhya Vihar, Borawan
Graduated 2025 · CGPA 7.34 / 10
Experience
3 roles since 2024
Software Engineer
Purple Sky Infotech
Building a real-time, multi-tenant voice-agent platform and the LLM infrastructure behind it.
- Architected a high-throughput, “true natural” voice-agent platform with Pipecat, orchestrating diverse STT/TTS engines and LLM APIs (OpenAI, Gemini) for sub-second conversational latency.
- Built a FastAPI backend managing multi-tenant agent sessions as a service, integrating Twilio & Plivo to scale concurrent active calls by 60%+.
- Deployed and optimized open-source LLMs on RunPod GPUs (RTX 3090) with vLLM — a 40% reduction in inference cost while holding real-time latencies.
- Built a RAG system (LangChain + Pinecone + OpenAI text-embedding-3-small) serving contextual answers over 100,000+ document chunks with sub-200ms query latency.
Core Contributor · Open Source
Sparrow — Techdome
Contributed reliability and response-handling features to an open-source API client.
- Added API response download and seamless image-response handling / UI rendering for REST APIs — 100% accurate binary image processing.
- Implemented WebSocket disconnection detection with real-time UI alerts, improving platform stability.
Software Developer Intern
Affimintus Technologies
Full-stack feature work focused on performance and authentication.
- Cut initial page load time by 25% and reduced average API response time by 20% for key user flows.
- Implemented JWT-based auth in Express and Recoil state management, improving scalability and reducing state-related bugs.
Stack
31 tools
Hover a tool to see where I've used it.
- AI / ML
- LLMs (OpenAI, Gemini)Used at Purple Sky Infotech, AssayRAGUsed at HootPR, AssayLangChainUsed at Purple Sky InfotechLangGraphUsed at HootPR, AssayvLLMUsed at Purple Sky InfotechPineconeUsed at Purple Sky InfotechPipecatUsed at Purple Sky InfotechPrompt Engineeringtext-embedding-3
- Backend
- FastAPIUsed at Purple Sky Infotech, HootPR, AssayExpressUsed at Affimintus Technologies, FeeloChatNode.jsUsed at Affimintus TechnologiesPythonUsed at Purple Sky Infotech, HootPR, AssayWebSocketUsed at Sparrow — TechdomeREST APIsUsed at Sparrow — TechdomeJWT Auth
- Data & Infra
- PostgreSQLUsed at HootPR, AssayMongoDBUsed at FeeloChatRedisUsed at HootPR, AssayAWS (EC2, S3, Lambda)DockerUsed at Purple Sky Infotech, HootPRRunPod GPUUsed at Purple Sky InfotechTwilioUsed at Purple Sky InfotechPlivoUsed at Purple Sky Infotech
- Frontend
- React.jsUsed at Sparrow — Techdome, Affimintus Technologies, FeeloChatTypeScriptUsed at Sparrow — Techdome, HootPR, Assay, FeeloChatTailwind CSSRecoilUsed at Affimintus Technologies, FeeloChatSocket.ioUsed at FeeloChat
- Languages
- JavaScriptTypeScriptUsed at Sparrow — Techdome, HootPR, Assay, FeeloChatPythonUsed at Purple Sky Infotech, HootPR, AssayC / C++
Projects
3 selected
- AI Agents
- DevTools
- Security
HootPR
AI code reviewer that reads, runs and fixes pull requests
An AI review platform for GitHub and GitLab. Every pull request is cloned into a sealed sandbox, mapped with a code graph, scanned by static analyzers and investigated by tool-using LLM agents — then a judge model throws out anything it can't back with evidence. What survives lands on the PR as a walkthrough, inline comments with one-click fixes and a merge check. Ask it, and it writes the fix, the tests or the CI repair too.
Agentic review pipeline
Triage → plan → investigator agents that read files, walk callers and callees, and run shell commands inside the sandbox — each capped at 8 tool calls and 60k tokens.
Built-in hallucination control
A finding must sit inside the diff, apply cleanly as a patch and pass an LLM judge's confidence bar. Linter output is evidence for the agents, never posted raw.
Sealed, disposable sandbox
Read-only, non-root, no capabilities, 768 MB; the network is cut after the clone and the container is destroyed after. Repo text is fenced off as untrusted input against prompt injection.
It writes the fix, too
@hootpr autofix, generate unit tests (run and iterated until they pass), fix CI from the failed job logs, resolve merge conflicts — pushed as a commit or a stacked PR.
Security suite
Blast radius traces changed code back to the HTTP routes it can reach; a repo-wide attack-surface map and an on-demand security architecture review.
Learns each team
Preferences taught in the PR thread are embedded in pgvector and recalled on the next review. Linked repos, remote MCP servers and web search add context.
Dashboard with a full trace
Every review's stages, agent steps, LLM calls, tool runs and judge verdicts; a Monaco diff workspace with findings inline; analytics, reports, audit log and a public REST API.
Measured, not guessed
An eval suite of seeded bugs, reversed CVE fixes and clean PRs scores precision, recall and cost per PR, with ablations that switch stages off.
- RAG
- Agents
- Finance
Assay
AI equity research agent — a report card for every stock
Type a US ticker and Assay reads the company's financials and SEC filings, computes every ratio in Python, benchmarks them against peers, and grades the business on profitability, financial health, growth, valuation and sentiment — then writes a cited investment memo with a BUY / HOLD / SELL call. The LLM writes the argument; it never produces a number and never picks the rating.
Every number is verified
Ratios are computed in pure Python. After the memo is written, every %, multiple and dollar figure is extracted and matched against those values — a transposed digit sends it back for a rewrite.
A rubric makes the call
27 metrics are scored 0–10 against named anchors and folded into five weighted dimensions. The model is told the rating and argues for it; if it disagrees, that goes in the caveats.
Every claim is cited
Qualitative sentences must cite a retrieved passage. Uncited claims, or citations to sources that don't exist, are removed before the memo is saved.
Retrieval led by the numbers
10-Ks streamed from SEC EDGAR plus news, chunked into pgvector. When revenue fell or leverage is high, it searches specifically for demand weakness or refinancing risk.
Peer benchmarking
Each metric gets a percentile against comparable companies — “operating margin 31%, better than 88% of peers” — with meaningless multiples filtered out.
“What would make it a BUY?”
The rubric runs backwards: a Scenario tab solves the smallest move in each metric that flips the rating, with draggable levers — no data fetch and no model call.
Live pipeline, full trace
Server-sent events stream each node as it finishes; the trace tab shows tokens, cost and citation rate for every run.
A complete product
Accounts with Google sign-in, compare, watchlist, PDF export, public share links, credits via Razorpay (test mode), quotas and a daily LLM budget guard.
- Realtime
- ML in browser
- E2E Encryption
FeeloChat
Secure, real-time, emotion-aware chat
An end-to-end encrypted chat app that shares real-time emotion. A lightweight ML model runs directly in the browser (face-api.js) to detect facial expressions, so raw video never leaves the device — privacy by design.
- In-browser facial-expression detection via face-api.js — zero server-side image processing.
- Socket.io realtime with sub-100ms processing and zero-lag sync across concurrent users.
- Complex UI state managed cleanly with Recoil.
Let's talk
Building something with AI, or need a backend that scales? Send me a note.
- Phone
- +91 9752588937
- Based in
- Indore, India
Or write here