The model to consider.
https://huggingface.co/thinkingmachines/Inkling
The closest model TensorRT-LLM already supports.
nemotron without hybrid attention, but interleaved SWA, Gemma3 style, MTP, nvfp4
What's your difficulty of supporting the model you want?
Just looking for support thats more performant than vllm/tokenspeed/sglang.
Before submitting a new issue...
The model to consider.
https://huggingface.co/thinkingmachines/Inkling
The closest model TensorRT-LLM already supports.
nemotron without hybrid attention, but interleaved SWA, Gemma3 style, MTP, nvfp4
What's your difficulty of supporting the model you want?
Just looking for support thats more performant than vllm/tokenspeed/sglang.
Before submitting a new issue...