Quick Start

Use the skills alongside your tests and benchmarks.

Prepare your environment

This example uses a local copy of tirx-kernels/tirx_kernels/basic/fp16_bf16_gemm.py on NVIDIA B200. It also needs a compatible CUDA-enabled PyTorch installation.

With the system prerequisites available, activate the Python environment your agent uses and install the harness:

git clone https://github.com/mlc-ai/TIRx-harness.git
cd TIRx-harness
git submodule update --init thirdparty/tvm-rust-ext
python -m pip install .

Check that its packages resolve:

python -c "import tvm.tirx, tvm_ffi, tirx_kernels, tirx_tools; print('Core imports OK')"

This checks imports only. See Installation for other installation methods.

Copy the skills to the directory your agent reads in TIRx-harness:

skills_dir="$PWD/.agents/skills"
mkdir -p "$skills_dir"
cp -R skills/tirx-wiki skills/tirx-debug-kernel skills/tirx-profile-kernel "$skills_dir/"
(cd "$skills_dir/tirx-wiki" && python scripts/fetch_references.py)

Use your agent’s discovery convention, such as .agents/skills or .claude/skills. Copy each skill as a complete directory. The fetcher downloads the wiki manuals and reference repositories; it requires network access.

Optimize a kernel

Clone tirx-kernels from the TIRx-harness repository root:

git clone https://github.com/mlc-ai/tirx-kernels.git

Start your agent (for example, Codex or Claude Code) from the TIRx-harness repository root.

Give your agent the following goal:

/goal Optimize a local copy of tirx-kernels/tirx_kernels/basic/fp16_bf16_gemm.py for FP16 C = A @ B.T, M=N=K=1024, on NVIDIA B200. Aim to match or beat torch-cublas while preserving correctness. Compare original, optimized, and cuBLAS timings on the same GPU with identical settings over three independent runs. Report latencies, speedups, and any remaining gap.

Read correctness before speedup. Compare original, optimized, and cuBLAS timings on the same GPU with identical settings over three independent runs. Check the reported latencies, speedups, and any remaining gap alongside the modified kernel.

Run an existing workload

Follow Optimization Runs for one complete example: install the setup dependencies, select the FP16 GEMM workload, prepare a worktree, launch an agent, and inspect the resulting frontier.