Article URL: https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request/
Comments URL: https://news.ycombinator.com/item?id=48321076
Points: 22
# Comments: 22
HA
Source
This article was originally published by Hacker News and written by NicoConstant.
Read original article on Hacker News