Skip to content
View ilyas-mallah's full-sized avatar

Block or report ilyas-mallah

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ilyas-mallah/README.md

Ilyas Mallah

I'm a software engineer working mostly in Luau, Rust, C++ and CUDA. My main interests are compilers and runtimes (JIT code generation, WebAssembly) and making code fast, from CPU hot paths to GPU kernels.

Recent work

  • llama.cpp #30077: decoding with a q4_0 KV cache ran at half speed on RTX 50 series GPUs in CUDA 12.8 builds. The fix makes long-context decoding up to 2.1x faster on an RTX 5090.
  • vLLM #60464: pick the SM90 block-FP8 CUTLASS kernel by batch size, 17 to 52% more serving throughput on an H100.
  • vLLM #60438: fall back from DeepGEMM instead of crashing at startup when nvcc is older than 12.9.
  • vLLM #60481: remove a residual copy per layer that a compiler pass left in block-FP8 models.

Earlier open-source work, mostly on the Luau compiler, its native code generator (x64 and AArch64) and its runtime, is under green-real, with seven changes merged into Luau in 2026.

More at ilyasm.dev. Open to internships and contract work: hello@ilyasm.dev

Popular repositories Loading

  1. llama.cpp llama.cpp Public

    Forked from ggml-org/llama.cpp

    LLM inference in C/C++

    C++

  2. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python

  3. ilyas-mallah ilyas-mallah Public

    Profile README

  4. wasmtime wasmtime Public

    Forked from bytecodealliance/wasmtime

    A lightweight WebAssembly runtime that is fast, secure, and standards-compliant

    Rust

  5. wabt wabt Public

    Forked from WebAssembly/wabt

    The WebAssembly Binary Toolkit

    C++

  6. llvm-project llvm-project Public

    Forked from llvm/llvm-project

    The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.

    LLVM