Research Focus

I work on making long-context and diffusion language models more efficient without giving up the behavior that makes them useful.

Inference Efficiency

Long-context LLM systems

KV-cache compression, sparse attention, distributed attention, and evaluation pipelines for scaling context length under real compute constraints.

Diffusion LLMs

Masking and feedback dynamics

Studying how remasking, interpolation, and embedding geometry affect convergence, hallucination lock-in, and fine-tuning stability.

Applied ML

Vision and multimodal systems

Computer vision, synthetic EO data, thermal imagery, retrieval systems, and civic intelligence dashboards.

Current Work

Research, systems, and applied ML that ship cleanly

I care about models that are not only interesting on paper, but also efficient, inspectable, and usable in real workflows.

Efficient inference KV-cache compression, sparse attention, distributed long-context systems
Diffusion LLMs Masking dynamics, interpolation behavior, and decoding stability
Applied ML builds Vision pipelines, research agents, civic dashboards, and deployment-ready tools

Experience

Research and engineering roles spanning long-context LLMs, distributed inference, and computer vision.

Shantanu Acharya, Research Scientist at NVIDIA

Research Intern
March 2026 – May 2026
  • Designed and implemented Pulsar Attention, a distributed long-context LLM inference framework that replaces static anchor duplication with content-aware statistical context summaries.
  • Built the experimental stack for distributed attention, Max-IDF summarization, FlashAttention-based inference, and evaluation on RULER and BABILong.
  • Co-authored a paper showing up to 4.7% accuracy improvement over dense attention, 6.1% over Star Attention, and 3.3x lower Phase-1 FLOPs.

Auric AI Labs

Computer Vision Intern
March 2026 – May 2026
  • Created an end-to-end synthetic EO data generation pipeline for object-detection models using controlled top-view renders from 3D military-object assets.
  • Developed a GPT-assisted semantic placement pipeline using probabilistic spatial priors to identify contextually valid object locations.
  • Combined BlenderProc rendering, Meshy AI asset generation, and RemoteCLIP-based semantic filtering to improve synthetic scene realism.

Selected Projects

Highlights from the larger project archive, now moved to a dedicated projects page for readability.

Civic ML

ParkWatch

AI-powered Bengaluru parking-enforcement intelligence dashboard with GraphSAGE hotspot forecasting, Mappls road-aware A* patrol planning, delay-exposure estimates, and exportable action plans.

Research Systems

Research Corpus Agent

Multi-agent scientific-literature RAG over 100K+ arXiv papers with hybrid retrieval, reranking, multimodal figure/table understanding, and evaluation pipelines.

Computer Vision

GATE-IR

Gated thermal perception pipeline routing infrared images through weather-specific enhancement before lightweight YOLO-based detection.

View all projects

Education

Indian Institute of Technology Roorkee

Bachelor of Technology in Electrical Engineering
Aug. 2024 – 2028 CGPA: 8.175/10 (3.538/4)*
  • * US scale conversion calculated via Scholaro methodology.

Recent Highlights

Selected milestones across research, hackathons, and campus activities.

  • 2nd Place, PixelPlay 2025 – campuswide computer vision hackathon.
  • Project selected for IIT Roorkee Institute Research Day 2025 through the TMI program.
  • Top 400 of 10,000 teams and selected for round 2 of Flipkart Gridlock 2.0.

View achievements and activities