View EnggTalha's full-sized avatar
🎯
Focusing
Talha Ahmed EnggTalha
🎯
Focusing
AI Engineer ||AI Agents || Generative AI || LLM || Livekit Developer || LLMOP ||Python
Stars
An OpenAI API compatible speech to text server for audio transcription and translations, aka. Whisper.
Voice agent using LiveKit (orchestration), Cartesia (STT + TTS), and OpenAI (LLM)
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
Official code repo for the O'Reilly Book - "Hands-On Large Language Models"