Portrait of Hemant Kumawat

Hemant Kumawat

Applied Scientist, Gaming AI @ Microsoft ยท Seattle, WA

I’m an Applied Scientist on Microsoft’s Gaming AI team, building real-time vision-language models that understand gameplay. My research asks how agents can learn world models โ€” compact enough to run in real time, rich enough to plan with, honest about what they can’t see โ€” and how to turn multimodal foundation models into agents that act, in games and on robots. Ph.D. from Georgia Tech with Prof. Saibal Mukhopadhyay; internships at Qualcomm, Amazon Robotics and CMU.

๐Ÿ”ฌResearch

Four open problems on the way to agents that model the world and act in it.

hover to scrub

World models

๐ŸŒHow do we learn world models an agent can plan with?

Generative models can imagine convincing futures, but planning needs models that respond to actions, stay consistent over long horizons and are cheap enough to run inside a control loop. I learn compact, action-conditioned dynamics โ€” Koopman embeddings that let a linear controller act directly from pixels, and object-centric 3D occupancy world models โ€” building toward controllable world models for games and robots.

illustrative

Multimodal foundation models

๐ŸŽฎHow do we turn foundation models that describe into agents that act?

Vision-language models can caption a frame, yet struggle to track state, ground language in action and stay truthful in real time. I work on closing that gap: real-time VLMs that understand gameplay at Microsoft and, in my Ph.D., decision transformers and Mamba policies that learn from interaction, and vision-action models adapted to new embodiments from little data.

drag an agent

Multi-agent world models

๐Ÿ•ธ๏ธHow should agents reason about what they can't see?

Real environments are partially observed: other agents are hidden, their interactions unknown, the future uncertain. I build stochastic generative models that infer hidden agents and their interactions from the ones we can see, spiking networks that learn interaction graphs from event streams, and forecasting for warehouse robot fleets โ€” toward multi-agent world models that plan with uncertainty rather than around it.

live sweep

Closed-loop perception

๐Ÿ“กHow can perception spend compute only where the task needs it?

Robots and on-device models can't afford to sense and process everything, everywhere. I build closed-loop perception in which the task decides what to look at: radar that steers a camera detector to the objects it misses (+14% recall, 3ร— less compute), LiDAR fired only where the camera points, and chirp-by-chirp radar with 3ร— lower latency โ€” ideas that carry over to adaptive compute in foundation models.

๐Ÿ“„Publications

15 papers ยท 161 citations ยท h-index 6 โ€” Google Scholar, Oct 2026

NameVenueYear TopicsRoleCites
๐Ÿ“„MAPLE: Multimodal Mamba Agent for Event Based Policy with Adaptive Value Estimation RA-L 2026 Robot learning First author 0
๐Ÿ“„DFDNet: Directional Feature Diffusion for Efficient Fully-Sparse LiDAR Object Detection Preprint 2025 Perception Co-author 0
๐Ÿ“„Toward Efficient and Robust Sequential Chirp-Based Data-Driven Radar Processing for Object Detection T-RS 2025 Perception Co-author 1
๐Ÿ“„LUGA: Lightweight Uncertainty-Guided Sensing Resolution Adaptation for Energy Efficient Radar Processing SENSORS 2025 PerceptionEdge AI Co-author 1
๐Ÿ“„Adaptive Graph Structure Inference for Learning Multivariate Point Processes using Spiking Neural Networks IJCNN 2025 Multi-agent Co-author 0
๐Ÿ“„Intelligent Sensing-to-Action for Robust Autonomy at the Edge: Opportunities and Challenges DATE 2025 Edge AI Co-author 21
๐Ÿ“„AdaCred: Adaptive Causal Decision Transformers with Feature Crediting AAMAS 2025 Robot learning First author 14
๐Ÿ“„RoboKoop: Efficient Control Conditioned Representations from Visual Input in Robotics using Koopman Operator CoRL 2024 Robot learningWorld models First author 12
๐Ÿ“„STEMFold: Stochastic Temporal Manifold for Multi-Agent Interactions in the Presence of Hidden Agents L4DC 2024 Multi-agentWorld models First author 7
๐Ÿ“„ChirpNet: Noise-Resilient Sequential Chirp Based Radar Processing for Object Detection IMS 2024 PerceptionEdge AI Co-first 5
๐Ÿ“„Cognitive Sensing for Energy-Efficient Edge Intelligence DATE 2024 Edge AI Co-author 2
๐Ÿ“„STAGE Net: Spatio-Temporal Attention-based Graph Encoding for Learning Multi-Agent Interactions in the Presence of Hidden Agents Preprint 2023 Multi-agentWorld models First author 4
๐Ÿ“„Radar Guided Dynamic Visual Attention for Resource-Efficient RGB Object Detection IJCNN 2022 Perception First author 18
๐Ÿ“„A Methodology for Understanding the Origins of False Negatives in DNN Based Object Detectors IJCNN 2022 Perception Co-author 5
๐Ÿ“„Task-Driven RGB-Lidar Fusion for Object Tracking in Resource-Efficient Autonomous System T-IV 2022 Perception Co-author 71

๐Ÿ’ผExperience

MS Applied Scientist, Gaming AI โ€” Microsoft 2026 โ€” now
  • Built efficient, real-time game-understanding VLMs for adaptive, context-aware game assistance, and persistent player profiles that power experiences such as game recaps and recommendations.
  • Designed evaluation pipelines for Microsoft Copilot that detect hallucinations and assess session and query quality, pinpointing failure points to accelerate targeted model and prompt improvements.

Vision-language-action modelsMultimodal learningEfficient adaptationPersonalizationModel evaluation

QC Research Intern โ€” Qualcomm ยท ADAS Vision May โ€” Aug 2025

Multi-modal, object-centric 3D occupancy world models ยท mentored by Amin Ansari

  • Designed a query-centric semantic occupancy framework using sparse, trackable 3D Gaussian queries anchored to object-level primitives for efficient scene understanding.
  • Proposed a Gaussian-to-voxel splatting strategy โ€” 2ร— faster runtime and 30% lower memory than projection-based baselines (in submission).

3D world modelsGaussian splattingNeRFDeformable attention

GT Graduate Research Assistant โ€” Georgia Tech ยท GREEN Lab Jan 2021 โ€” Dec 2025

advised by Saibal Mukhopadhyay

  • Efficient RL โ€” sample-efficient multimodal RL with contrastive spectral Koopman encoding: 10ร— lower compute and +20% accuracy (CoRL 2024).
  • Long-horizon sequence learning โ€” adaptive causal decision transformers with separate local and long-horizon representations for offline RL (AAMAS 2025).
  • Multi-agent dynamics โ€” stochastic generative graph models with neural ODEs and spatiotemporal graph attention for partially observable systems (L4DC 2024).
  • Closed-loop perception โ€” noise-resilient RGB / LiDAR / radar processing that adapts memory and compute to real-time needs (IEEE T-IV, IJCNN, IMS).
  • Vision-action adaptation โ€” guiding a vision-action model’s generation toward a target domain for cross-embodiment and cross-task transfer with limited data.

Causal decision transformersOffline RLContrastive learningMambaNeural ODEsSensor fusion

AR Applied Scientist II Intern โ€” Amazon Robotics May โ€” Aug 2024

Multi-agent probabilistic behavior models for robots ยท mentored by Andreas Kolling

  • Designed a goal-conditioned forecasting model to predict slowdowns and deadlocks of Proteus robots in dense warehouses.
  • Built a scalable multi-agent interaction engine estimating inter-robot influence across 20 TB of real-world sensing and occupancy-grid data.
  • Integrated goal-aware trajectory prediction with planning to reduce navigation conflicts โ€” an order-of-magnitude improvement in coordination.

Multi-agent forecastingGoal-conditioned planningMotion prediction

RI Research Scholar โ€” Robotics Institute ยท Carnegie Mellon May โ€” Aug 2019

Planning & control for evasive maneuvers in autonomous vehicles ยท mentored by John M. Dolan

  • Combined iLQR with RRT* to plan evasive maneuvers with large tire slip โ€” including drifting โ€” and deployed the planner in ROS simulation and on an RC car avoiding suddenly-appearing obstacles at speed. (code)

Motion planningiLQRROS

IITB Student Project Lead, Self-Driving Car Team โ€” SeDriCa ยท UMIC, IIT Bombay 2019 โ€” 2020
  • Led an interdisciplinary team of 50 students across three international robotics competitions.
  • Managed an INR 4.5M budget and built a vendor & sponsorship network worth INR 3M with Velodyne, Ouster, Continental, Aptiv and NVIDIA.

LeadershipFull-stack autonomy

๐ŸŽ“ Education

  • Georgia Institute of Technology โ€” M.S. & Ph.D., Electrical & Computer Engineering ยท 2021 โ€” 2025 ยท Advisor โ€” Prof. Saibal Mukhopadhyay
  • Indian Institute of Technology Bombay โ€” B.Tech, Mechanical Engineering & Computer Science ยท 2016 โ€” 2020

๐Ÿ“ฐNews

2026Joined the Gaming AI team at Microsoft in Seattle as an Applied Scientist.
2026MAPLE, a multimodal state-space Mamba agent for event-based policies, accepted to IEEE Robotics and Automation Letters.
2025Completed my Ph.D. in ECE at Georgia Tech. ๐ŸŽ“
Oct 2025LUGA at IEEE SENSORS 2025 and sequential-chirp radar processing in IEEE Transactions on Radar Systems.
May 2025Research internship at Qualcomm ADAS on object-centric 3D occupancy world models with Gaussian queries.
Dec 2024AdaCred โ€” adaptive causal decision transformers with feature crediting โ€” accepted at AAMAS 2025. ๐ŸŽ‰
Nov 2024Presenting RoboKoop at CoRL 2024 in Munich.
Sep 2024RoboKoop accepted at CoRL 2024. ๐ŸŽ‰
Jul 2024Presenting STEMFold at L4DC 2024 in Oxford, UK.
May 2024Applied Scientist II internship at Amazon Robotics on multi-agent behavior forecasting.

๐ŸŽคService & talks

Reviewer

NeurIPS 2024 & 2025ICLR 2023 & 2025CoRL 2025AAAI 2025AISTATS 2025IJCNN 2022 & 2024

Talks

  • Vision-Language Models 101
    Cohere For AI ยท community talks with Vaishaal Shankar
  • Task-Driven Model Learning
    Cohere For AI

Mentoring

  • Georgia Tech SURE โ€“ Intel REU
    2023 ยท ROS sensing & perception stack for the Unitree A1 quadruped
  • Ph.D. student mentor, Georgia Tech
    2022 โ€“ 2025 ยท analog-to-feature object detection, event-camera activity recognition, CARLA closed-loop control
  • Autonomous Robotics Summer Program, IIT Bombay
    2018 โ€“ 2019 ยท weekly lectures on localization, vision, planning & control for ~40 students
๐Ÿ“ฌ

Get in touch. Happy to talk research, collaborations or ideas โ€” email is fastest.

Last edited October 2026

โ†‘โ†“ to navigateโ†ต to openesc to close