Hardware + AI engineer focused on model-hardware co-design to deliver high-throughput, energy-efficient inference.
Highlights:
- π§ͺ Binary Neural Networks (BNN): 1-bit models to reduce compute and memory while retaining performance.
- π§ High-performance model optimization: operator fusion and hardware/compiler tuning.
- β‘οΈ LLM inference acceleration: quantization, kernel and operator optimization, and parallel deployment.
π Skills
Python C++ HLS Verilog CUDA Triton