Long-context LLM systems
KV-cache compression, sparse attention, distributed attention, and evaluation pipelines for scaling context length under real compute constraints.
I work on making long-context and diffusion language models more efficient without giving up the behavior that makes them useful.
KV-cache compression, sparse attention, distributed attention, and evaluation pipelines for scaling context length under real compute constraints.
Studying how remasking, interpolation, and embedding geometry affect convergence, hallucination lock-in, and fine-tuning stability.
Computer vision, synthetic EO data, thermal imagery, retrieval systems, and civic intelligence dashboards.
I care about models that are not only interesting on paper, but also efficient, inspectable, and usable in real workflows.
A small selection of recent work. The full publication list has links and BibTeX entries.
Research and engineering roles spanning long-context LLMs, distributed inference, and computer vision.
Highlights from the larger project archive, now moved to a dedicated projects page for readability.
AI-powered Bengaluru parking-enforcement intelligence dashboard with GraphSAGE hotspot forecasting, Mappls road-aware A* patrol planning, delay-exposure estimates, and exportable action plans.
Multi-agent scientific-literature RAG over 100K+ arXiv papers with hybrid retrieval, reranking, multimodal figure/table understanding, and evaluation pipelines.
Gated thermal perception pipeline routing infrared images through weather-specific enhancement before lightweight YOLO-based detection.
Selected milestones across research, hackathons, and campus activities.