C++ / CUDA / LLM inference systems · Maintainer of @open-infra-ai
- shenzhen
-
16:28
(UTC +08:00) - https://github.com/open-infra-ai
Pinned Loading
-
open-infra-ai/paged-serving
open-infra-ai/paged-serving PublicRust LLM Serving 控制面:Paged KV、continuous batching、OpenAI SSE、C ABI 后端与可信压测
Rust 1
-
open-infra-ai/cuflash
open-infra-ai/cuflash Publiccuflash — FlashAttention from scratch in CUDA C++: scalar → WMMA tensor-core forward, FlashDecoding/Split-KV, Roofline analysis. FP16/BF16/FP32 fwd/bwd, sm_70–sm_90.
Cuda
-
open-infra-ai/tiny-llm
open-infra-ai/tiny-llm PublicCUDA/C++ LLM 推理运行时:GGUF、W8A16、tokenizer、Paged KV、CUDA Graph 与可复现基准
C++
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.




