Universal LLM Deployment Engine with ML Compilation
Role in this project:
Back-end Developer & ML Engineer Contributions:22 reviews, 23 PRs, 6 pushes in 1 year 6 months
Contributions summary:Yixin contributed to the memory optimization of the model building process by modifying the Llama attention mechanism and the causal mask generation in the `mlc_llm/relax_model/llama.py` file. They also worked on the implementation of a BNF AST and parser for EBNF grammar, introducing new code for parsing and simplifying grammars. Additionally, the user integrated a JSON grammar into the generation pipeline.
language-modelllmmachine-learning-compilationtvm
Open deep learning compiler stack for cpu, gpu and specialized accelerators
Contributions:155 pushes, 56 branches in 1 year 3 months