ML engineer, inference · [company]
[2024 - present]
- [what got faster, by how much, on which hardware]
- [what you built or owned to get there]
ML systems engineer: GPU inference, low-precision formats, systems programming.
coldblime.com · github.com/1morello
[2024 - present]
[2021 - 2024]
[degree] · [university]
[years]
fp4-speech. Hardware FP4 inference for speech on Blackwell: FP4 weights safe across three architectures, 2x Orpheus decode with CUDA graphs, activation failure modes measured and mapped.
in progress