Generative AI, computer vision, and robot learning banner
AI Research Engineer | Generative AI | Computer Vision | Robot Learning
I build research-driven machine learning systems, from model design and large-scale training to efficient production deployment.
LinkedIn Google Scholar IsoFM on OpenReview Email
I build research-driven machine learning systems that move from model design and distributed training to optimized production inference.
I am an AI Research Engineer with 4+ years of experience across generative AI, computer vision, multimodal learning, and efficient inference. At Fynd, I develop image and video systems involving Flow Matching, diffusion models, super-resolution, model distillation, distributed PyTorch, and TensorRT.
Previously, at Wobot.ai, I built large-scale video analytics systems for detection, multi-object tracking, and cross-camera association. My current research connects generative vision with embodied intelligence: learning representations, trajectories, and policies that are accurate, efficient, and deployable.
4+ years of experience 20K daily active users 1,000+ cameras 95.5% robot-learning success rate
Accepted at the ICML 2026 Workshop on Structured Probabilistic Inference and Generative Modeling (SPIGM).
A geometry-aware Flow Matching formulation for straighter generative transport paths and more efficient generation.
Sep 2023 - Present · Flow Matching super-resolution, single-step distillation, SDXL and FLUX training, video segmentation and inpainting, controllable generation, distributed PyTorch, and TensorRT deployment.
Feb 2022 - Sep 2023 · Production detection and tracking, multi-camera analytics, CPU and GPU pipeline optimization, Docker, and NVIDIA Triton deployment.
🤖 mini-pi0
Multimodal robot-learning framework using ManiSkill, MuJoCo, wrist-camera observations, and visuomotor Flow Matching policies.
95.5% StackCube-v1 success · VLA · IL · RL
View benchmarks →Conditional Flow Matching experiments for image and video generation with optimal-transport paths and ODE sampling.
13+ stars · PyTorch · Generative Modeling
Complete video-generation training and inference pipeline using MeanFlow with a Diffusion Transformer.
Video Generation · DiT · Flow Models
🌙 Zero-DCE
TensorFlow implementation of zero-reference low-light image enhancement.
49+ stars · 8+ forks
More applied ML work: 🩺 PraNet Polyp Segmentation · 🎙️ TinyML Audio Classification
I am building an SO-ARM101 from individual components rather than using a pre-assembled system. The goal is to connect simulation-based policy development with physical data collection, evaluation, and learned-policy deployment. Hardware integration and real-world policy testing are currently in progress.
Python PyTorch TensorFlow OpenCV Linux
Representation Learning Multimodal Learning Distributed Training DDP FSDP
Diffusion Models Flow Matching SDXL FLUX Super-Resolution Flow Distillation One-Step Generation Image Generation Video Generation IP-Adapters Textual Inversion LoRA PEFT
Object Detection Multi-Object Tracking Image Segmentation Video Segmentation Image Restoration Video Restoration Vision-Language Models Controllable Generation
MuJoCo NVIDIA Isaac Sim LeRobot
Vision-Language-Action Models Robotic Manipulation Visuomotor Policies Imitation Learning Reinforcement Learning Flow Matching Policies ManiSkill Robosuite Sim2Real
NVIDIA Docker FastAPI Git Weights & Biases
NVIDIA Triton Inference Server MLflow GPU Optimization Model Serving Data Curation Experiment Tracking
Bachelor of Science in Computer Science, University of Mumbai, 2022
CGPA: 9.21/10
I am interested in research and engineering conversations around generative vision, efficient deep learning, multimodal intelligence, and learning-based robotics.
- Email: mail2tauhidkhan@gmail.com
- LinkedIn: tauhid-khan-24bb45177
- Google Scholar: Tauhid Khan
- OpenReview: Isokinetic Flow Matching
Research depth. Engineering rigor. Systems that run outside the notebook.