Skip to content

Navigation Menu

Sign in
Sign up
@zhenye234
zhenye234
Follow

YE Zhen zhenye234

🍉
Speech synthesis, Audio generation, Speech LLM

Block or report zhenye234

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official implementation for paper "Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction". We treat is as a new decoder for Music Reconstruction.

Python 73 2 Updated Sep 1, 2026
Python 8,186 559 Updated Aug 15, 2026

ACM MM 2026 Talker-T2AV Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

Python 86 3 Updated Aug 1, 2026

LIA-X: Interpretable Latent Portrait Animator

Python 106 12 Updated Sep 17, 2025

The official repo for SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.

Python 44 2 Updated May 26, 2026

MiMo-Audio: Audio Language Models are Few-Shot Learners

Python 1,080 107 Updated Jun 17, 2026

Open-Source Frontier Voice AI

Python 53,843 6,081 Updated Sep 3, 2026

Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.

Python 3,020 314 Updated Jan 26, 2026

Text-audio foundation model from Boson AI

Python 8,346 640 Updated Jun 5, 2026
Python 303 39 Updated Jul 22, 2025

This is the code for paper: XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Python 97 5 Updated Sep 19, 2025

Llasa Speed Up

Python 65 6 Updated Jan 18, 2026
Python 133 12 Updated Sep 1, 2026

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation

Python 4,734 375 Updated Jun 21, 2025
Python 6,095 475 Updated Jun 15, 2026

LLaSA WebUI using ExLlamaV2 and FastAPI.

Python 28 5 Updated Mar 30, 2025
Python 350 43 Updated Apr 11, 2025

Official repository of the paper "MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization".

Python 372 22 Updated Aug 4, 2025

MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Python 770 30 Updated Sep 7, 2025

LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement

Python 106 19 Updated Apr 1, 2025

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Python 222 15 Updated Feb 28, 2025

Explore the Multimodal "Aha Moment" on 2B Model

Python 623 23 Updated Mar 18, 2025
Python 170 9 Updated Nov 22, 2024

[ICML 2025] SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation

Python 318 33 Updated Nov 5, 2025

Official implementation of the source-filter HiFiGAN vocoder

Python 278 35 Updated Jul 29, 2023

Zonos-v0.1 is a leading open-weight text-to-speech model trained on more than 200k hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS ...

Python 7,244 816 Updated Mar 5, 2025
Next

AltStyle によって変換されたページ (->オリジナル) /