tirx-lite

Kernel DSL & Kernel Zoo

A PTX-level subset of the TIRx foundation IR, with a traced Python authoring API.

Key idea

TIRx is the foundation IR; tirx-lite is a subset of it for kernel programming. A tirx-lite kernel is represented as a TIRx PrimFunc and uses the TIRx compiler and analysis infrastructure. The traced Python DSL is the authoring interface to this subset.

Its defining feature is PTX-level abstraction: instruction operations, operands, memory addresses, synchronization, and execution roles are explicit in the kernel source. This gives an agent concrete operations to edit and inspect when improving a kernel.

What the subset includes

Part

Included in tirx-lite

Values and local storage

Typed scalar expressions, arithmetic, casts, and register-local scalars and arrays.

Control flow

Loops, conditional branches, and instruction predication.

Memory and addressing

Typed global pointers, shared-memory allocation, address calculation, and instruction operands such as tensor maps and shared-memory descriptors.

Instruction operations

PTX operations for memory movement, computation, synchronization, and matrix instructions, plus supported CUDA intrinsics.

Execution structure

Kernel launch configuration, CTA/warp/lane/thread coordinates, and explicit warp roles.

The authoring API also supplies barrier and pipeline helpers, stage/phase cursors, and reusable instruction sequences. These helpers compose the low-level operations in the subset.

The subset excludes high-level tile primitives and general tensor-layout APIs. Ordinary tensors use the default layout; global and shared-memory data accesses use explicit instructions. Shared-memory swizzles remain explicit allocation options.

Core API

Import the Python interface as tirx_kernels.tirx_lite, conventionally named txl. Decorating a function with @txl.kernel(...) traces it and returns a Kernel object.

Interface

Role

txl.kernel(...)

Define a kernel entry and its launch configuration.

txl.gptr(dtype), txl.TensorMap, scalar dtype annotations

Describe pointer, tensor-map, and scalar parameters.

txl.local_scalar(...), txl.alloc_local(...), txl.assign(...)

Allocate and update register-local values.

txl.If(...), txl.While(...), txl.serial(...), txl.unroll(...)

Build IR branches and loops through context managers.

txl.cta_id(...), txl.warp_id(), txl.lane_id(), txl.thread_id()

Obtain execution coordinates for addressing and control flow.

txl.smem_pool().alloc(...)

Allocate shared-memory storage with an explicit shape, dtype, and optional swizzle.

txl.specialize()

Define named warp roles and their execution scopes.

txl.MBarrier, txl.TMABar, txl.TCGen05Bar

Construct barrier helpers.

txl.Pipeline, txl.PipelineState, txl.RingState

Construct pipeline protocols and stage/phase state.

txl.ptx, txl.cuda, txl.idioms

Spell instructions and CUDA intrinsics, or compose reusable instruction sequences.

kernel.func / kernel.mod

Inspect the pre-lowering TIRx PrimFunc or its containing IRModule.

kernel.compile(target=None)

Compile through the TIRx pipeline.

The authoring guide owns the full language contract. For exact signatures and supported forms, use the entry API source and the public namespace matching your installed revision. The default kernel dependency in pyproject.toml tracks the upstream default branch.

Use it on its own

Install the Python environment. tirx-lite is distributed in tirx-kernels; importing it does not require an agent or kcoral.

Save this as zero.py. It constructs a kernel that writes one float per lane to a 32-element output buffer:

import tirx_kernels.tirx_lite as txl


@txl.kernel(warps=1, arch="sm_100a", grid=1)
def zero(out: txl.gptr(txl.f32)):
    txl.ptx.st.global_.f32(out.ptr_to([txl.lane_id()]), txl.float32(0))


print(zero.func)
python zero.py

The result is a printed TIRx function, not a GPU launch. Compiling with zero.compile() additionally needs a compatible CUDA toolchain; running the result needs a GPU supporting the kernel’s target.

Note

Keep kernel definitions in a Python file and use live type annotations. Do not enable from __future__ import annotations in a tirx-lite module.

Kernel Zoo

The tirx-kernels collection provides complete kernels with input preparation, correctness checks, and benchmark entry points. Use it to find a nearby implementation before writing a new kernel.

Start with

Examples

Basic kernels

FP16/BF16 GEMM and RMSNorm.

Attention and library ports

FlashAttention, FlashInfer, and cuDNN.

Curated native kernels

Maintained implementations and measured results.

List the kernels and configurations available in your installed package:

python -m tirx_kernels.registry --format json

For example, run a correctness check of the GEMM implementation:

from tirx_kernels.basic import fp16_bf16_gemm

fp16_bf16_gemm.run_test(dtype="fp16", M=1024, N=1024, K=1024)

This compiles and runs on a GPU supported by that kernel. Check its KERNEL_META for supported architectures and its source for dependencies and configuration choices. The package’s catalog and benchmark results describe the implementations at that revision.

To add a kernel produced by your agent, follow Contribute a kernel.

Connect it to your agent loop

Use .func with Synccheck and Racecheck. Bind every runtime argument to concrete inputs for the invocation you want to inspect. Compile the kernel separately and give its launch callable to your test or benchmark adapter.

Keep compile and setup work outside the timed callable when the task measures GPU execution latency. Each benchmark owns its own candidate interface; a Kernel object alone does not define the workload’s input or output contract.

The wiki skill helps an agent find the nearest implementation in the wiki’s fetched references/repos/tirx-kernels/ checkout. Verify runtime APIs against the installed package.

Use another kernel language

Connect your language’s compiler and launch interface to the workload. You can keep tirx_tools if the resulting TIRx uses operations it supports; another IR needs corresponding analysis tools. Supply examples for the new language.

Keep the workload’s input domain, correctness policy, and scoring rule fixed while changing components. For packaged tasks, the task declaration also defines the permitted kernel language. To define a different contract, see Add a workload.