YE Zhen zhenye234
-
Hong Kong University of Science and Technology
- Hong Kong
- @zhenye234
- https://huggingface.co/ZhenYe234
- in/zhen-ye-25734a358
Stars
Official implementation for paper "Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction". We treat is as a new decoder for Music Reconstruction.
ACM MM 2026 Talker-T2AV Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
The official repo for SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.
MiMo-Audio: Audio Language Models are Few-Shot Learners
Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.
Text-audio foundation model from Boson AI
This is the code for paper: XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
Official repository of the paper "MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization".
MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
Explore the Multimodal "Aha Moment" on 2B Model
[ICML 2025] SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
Official implementation of the source-filter HiFiGAN vocoder
Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS ...