Joined August 2018
- How Profitable is LLM Inference? Doing the Math on Kimi K3
- Exploring Speculative Decoding
- Deep dive into ideas from the excellent @vllm_project: KV cache, paged attention, tensor parallelism.
Get the full app experience
Unlock more features and see what people are talking about right now.

