
Birat Gautam
AI/ML Engineer @ NepaWorks
CS with AI @ Birmingham City University
Currently working with LLMs, agentic AI, RAG, and NLP — turning research into systems that ship.
Curiosity drives innovation. I like understanding how things work under the hood, learning to explain, and building to contribute.
11 followers · 69 repos
ॐ पूर्णमदः पूर्णमिदं पूर्णात्पूर्णमुदच्यते ।
पूर्णस्य पूर्णमादाय पूर्णमेवावशिष्यते ॥
Activity
GitHub contributions over the years.
Experience & Education
NepaWorksAI/ML Engineer- Architected an end-to-end multimodal lead-qualification platform processing 300,000+ commercial retail businesses and 65,000+ store photos across distributed Playwright scrapers, local LLM inference, and human review, achieving 99.95% reliability.
- Serve Gemma 4 on-prem on an RTX 5090 via a Blackwell-tuned llama.cpp build (300-400 classifications/hour).
- Achieved 99.0% to 100% lead precision (0 false leads on held-out benchmarks) and 88% product-line accuracy at $0 marginal API cost (matching 97% of cloud Gemini Pro/Flash reference accuracy) by upgrading on-prem inference from 12B to Gemma 4 31B Dense.
- Traced classification errors to missing image detail rather than the prompt by testing one change at a time; raised image resolution to 2048px (1,120 vision tokens/image, up to 15 photos / about 18.5K tokens per store), which fixed blurred shelf-level products.
- Eliminated VLM confidence miscalibration (where 90% of false negatives reported 1.00 confidence) by replacing score thresholds with JSON-schema grammar citations requiring exact photo indexing, explicit verification flags, and a deterministic 2% live audit sample.
- Built the async FastAPI and PostgreSQL control plane for distributed workers: lease-based claiming, transaction-scoped advisory-lock leader election, PgBouncer-safe; 35 migrations and 2,400+ automated tests.
- Optimized the pipeline: cut export latency 24x (448 ms to 18 ms) via batched queries, reduced host server RAM 75% (36 GB to 9 GB) via sliding-window checkpoint caps (--ctx-checkpoints 4), and cut CDN download bandwidth by about 90%.
Synapse TechnologiesJr. AI/ML Engineer- Engineered a multilingual NLP pipeline for English, Nepali, and Romanized Nepali queries using transformer-based intent classification and script-aware NER; reduced inference latency from ~10s to ~30ms via ONNX optimization and INT8 quantization.
- Developed and productionized the KanoonBox RAG pipeline using ChromaDB vector search and NepSearch-based legal grounding; optimized intent-routing architecture to reduce reasoning latency from ~5s to ~25ms while maintaining response quality.
- Deployed Nepali document OCR using PaddleOCR-VL-0.9B via llama.cpp within 2 GB VRAM (7s/page); integrated Qwen2.5-VL-3B-Instruct-GGUF into Synapse Intelligence for enterprise document understanding and knowledge delivery.
- Built Nepal's first structured Nepali ASR corpus from YouTube audio using VAD, NFC normalization, and Gemini validation; LoRA fine-tuned Whisper-large-v3-turbo on resource-limited hardware (WER: 33, WandB monitored).
- Benchmarked lightweight Gemma 4 variants on CPU infrastructure using Ollama, llama.cpp, and GGUF quantization formats for on-device LLM deployment feasibility.
Birmingham City UniversityBSc (Hons) CS with AI- Currently Third Semester Student.
- Spearheaded a team of five to build a face-recognition attendance web app.
- Completed First Year with First Class Honours.
Deerwalk Compware Ltd.Backend Developer Intern- Designed and developed an Attendance Management System using Django.
- Implemented Excel report generation for detailed attendance records.
- Built an Admin Panel with full CRUD operations.
Uniglobe SS/CollegeHigh School Graduate- Graduated with a 3.78 / 4 GPA.
- Awarded Student of the Year — UGSS Achievers Award 2022.
Research
Automated Nuclei Detection in Microscopy Images
International Journal on Engineering Technology (InJET)
Developed a U-Net-based deep learning model for nuclei detection in microscopy images (2018 Data Science Bowl dataset), achieving IoU 0.88. Automates segmentation, improves accuracy over manual annotation, and enhances diagnostic capabilities in resource-constrained medical environments.
Projects
Atri — Local Agentic Coding CLI
A local-first agentic coding CLI powered by Gemma 4 via llama.cpp — designed to match Claude Code's capabilities without sending code to the cloud. Fe…
EPL Match Prediction Platform
Fully local, containerized data and ML platform for English Premier League match prediction. Apache Airflow orchestrates data ingestion, validation, f…
Screen Recording & Video Sharing Platform
Full-stack screen recording and video sharing platform. Users record their screen in the browser, upload to Bunny Stream/Storage CDN, and control publ…
Automated Nuclei Segmentation
U-Net deep learning model for automated nuclei detection in microscopy images, benchmarked on the 2018 Data Science Bowl dataset with an IoU score of …
Real-Time Chat Application
Full-stack chat app with real-time group and private messaging via Socket.IO. Backend: Express + TypeScript, MongoDB for message persistence, JWT auth…
WebRTC × WebSockets P2P Video
Research and implementation workspace for browser-to-browser peer-to-peer WebRTC video, using a Django Channels WebSocket server as the signaling laye…
Attendance System — Face Recognition
Django-based attendance management system with real-time face recognition powered by PyTorch and WebSocket video streaming. Teachers manage courses an…




