Writing on inference, LLMs and the future of AI infrastructure.
Training got all the attention. Now the hard part is serving.
How we got model spin-up down from 30 seconds to 8ms.
A practical walkthrough of the fastest model iteration loop.