GitHub Org Non-Profit License Active
Towards Real-World Perception and Modeling โ
building open AI systems that perceive, understand, generate, and act.
UniX-AI-Lab is an open, non-profit research collective working toward:
We build intelligent systems that can continuously perceive, understand, model, and interact with the real world.
Our research spans four closely connected directions:
Multimodal Perception & Understanding ยท Visual Generation & Editing
Unified World Modeling ยท Agentic Reasoning
We are particularly interested in Multimodal Large Language Models (MLLMs) for video understanding, streaming perception, long-context multimodal reasoning, and real-world interaction.
We believe meaningful progress requires connecting perception and generation: an intelligent system should not only recognize what is happening, but also understand temporal dynamics, predict future states, reason about actions, and model how the world evolves.
Our work combines academic rigor with practical capability. We care deeply about why things work, not only that they do. Every model, dataset, benchmark, and codebase we produce is released openly whenever possible.
Building multimodal systems that understand real-world visual and temporal information.
Topics we explore:
- Multimodal Large Language Models (MLLMs)
- Image and video understanding
- Streaming and online video understanding
- Long-video and long-context reasoning
- Temporal grounding and event understanding
- Audio-visual-language alignment
- Embodied and interactive perception
Developing generative models that create, edit, and simulate realistic visual worlds.
Topics we explore:
- Image, video, and 3D generation
- Diffusion and flow-based models
- Instruction-following visual editing
- Controllable and compositional generation
- Identity, layout, motion, and style control
- 3D-consistent generation
- Reinforcement learning for generative models
Unifying perception, generation, and prediction to model how real-world states evolve.
Topics we explore:
- Unified understanding and generation
- World models and future-state prediction
- Any-to-any multimodal architectures
- Physical, social, and causal reasoning
- Multimodal tokenization and representation learning
- Cross-modal grounding and alignment
- Interactive environment modeling
Building systems that reason, plan, use tools, and act in dynamic environments.
Topics we explore:
- Multimodal autonomous agents
- Tool-augmented reasoning
- Long-horizon planning and decision-making
- Multi-agent collaboration
- Reinforcement learning for reasoning
- Scientific discovery automation
- Agents grounded in streaming perception
๐ WorldReasonBench
Human-aligned stress testing of video generators as future world-state predictors.
WorldReasonBench reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose evolution remains physically, socially, logically, and informationally consistent?
๐ฌ 436 curated cases ยท 4 reasoning dimensions ยท 22 subcategories
๐ค 11 closed- and open-source generators benchmarked head-to-head
๐งโโ๏ธ WorldRewardBench: ~6K expert preference pairs over 1.4K videos
๐ ScorePR โ human ฯ = 0.955
Video Generation World Model Benchmark Reward Model
๐ฌ StreamOPD
A post-training recipe with spatio-temporal cue gating for streaming video understanding.
StreamOPD fixes a deliberately austere inference path โ four recent frames, no memory bank, no retrieval, no reasoning trace โ and asks how far post-training alone can go. ST-CueGate then gates on-policy distillation by how much a grounded visual cue actually shifted the teacher's likelihood on the student's own response.
๐ฏ 84.6 StreamingBench ยท 69.3 OVO-Bench macro โ a 4B student past its own 9B teacher ๐๏ธ 4 recent frames at 1 fps ยท memory-free, retrieval-free inference ๐งช 25K verifiable video QA items with the full construction pipeline released ๐ค Weights on the Hub, evaluable straight from the repo id
Streaming Video MLLM On-Policy Distillation Post-Training
We are actively developing projects toward real-world perception and modeling.
In the pipeline:
- Long-video perception and temporal reasoning benchmarks
- Unified vision-language understanding and generation
- Instruction-driven video generation and editing
- Interactive world models for future-state prediction
- Multimodal agents for real-world reasoning
Watch this space โ or better yet, join us.
Open Research โ Slow Research
Non-Profit โ Low Quality
Accessible โ Trivial
We operate by a few core principles:
| Principle | What it means in practice |
|---|---|
| Radical openness | Code, weights, data, and evaluations โ public whenever possible |
| Depth over breadth | We would rather understand one thing deeply than five things superficially |
| Community first | Research should be reproducible, readable, and useful to everyone |
| No gatekeeping | Research direction should be guided by scientific value and community impact |
| Frontier, not incremental | We pursue fundamental contributions rather than marginal benchmark gains |
| Perception meets modeling | Understanding the real world requires connecting observation, prediction, generation, and action |
Publication list coming soon. All work will be posted to arXiv and linked here.
We target top-tier venues including NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, and EMNLP, and release preprints as soon as they are ready.
UniX-AI-Lab is open to researchers, engineers, and students who share our mission.
Ways to contribute:
- ๐ Issues & PRs โ bug fixes, new features, and documentation improvements
- ๐ฌ Research discussions โ open an issue to discuss ideas before implementing
- ๐ Writing โ help us write papers, blog posts, or tutorials
- ๐ Reproduction โ replicate prior work and share what you find
- โญ Star โ help other researchers discover our work
We especially welcome:
- ML researchers at any career stage, including PhD students
- Engineers who care about open-source research infrastructure
- Researchers interested in multimodal perception, video understanding, generation, world models, and agents
- Anyone passionate about responsible and transparent AI development
| ๐ GitHub | github.com/UniX-AI-Lab |
| ๐ง Email | Open an issue โ we respond to all of them |
| ๐ Website | Coming soon |
We have no corporate sponsors and no commercial agenda โ just curiosity and open science.
If you use our work, please cite it. If you improve it, please share it.