Skip to content

[New Model]: thinkingmachines/Inkling #16507

Description

@michaelfeil

The model to consider.

https://huggingface.co/thinkingmachines/Inkling

The closest model TensorRT-LLM already supports.

nemotron without hybrid attention, but interleaved SWA, Gemma3 style, MTP, nvfp4

What's your difficulty of supporting the model you want?

Just looking for support thats more performant than vllm/tokenspeed/sglang.

Before submitting a new issue...

  • Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.

Metadata

Metadata

Assignees

Labels

new modelRequest to add a new model

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions