Quick Start¶
Use the skills alongside your tests and benchmarks.
Prepare your environment¶
This example uses a local copy of
tirx-kernels/tirx_kernels/basic/fp16_bf16_gemm.py on NVIDIA B200. It also needs
a compatible CUDA-enabled PyTorch installation.
With the system prerequisites available, activate the Python environment your agent uses and install the harness:
git clone https://github.com/mlc-ai/TIRx-harness.git
cd TIRx-harness
git submodule update --init thirdparty/tvm-rust-ext
python -m pip install .
Check that its packages resolve:
python -c "import tvm.tirx, tvm_ffi, tirx_kernels, tirx_tools; print('Core imports OK')"
This checks imports only. See Installation for other installation methods.
Copy the skills to the directory your agent reads in TIRx-harness:
skills_dir="$PWD/.agents/skills"
mkdir -p "$skills_dir"
cp -R skills/tirx-wiki skills/tirx-debug-kernel skills/tirx-profile-kernel "$skills_dir/"
(cd "$skills_dir/tirx-wiki" && python scripts/fetch_references.py)
Use your agent’s discovery convention, such as .agents/skills or
.claude/skills. Copy each skill as a complete directory. The fetcher downloads
the wiki manuals and reference repositories; it requires network access.
Optimize a kernel¶
Clone tirx-kernels from the TIRx-harness repository root:
git clone https://github.com/mlc-ai/tirx-kernels.git
Start your agent (for example, Codex or Claude Code) from the
TIRx-harness repository root.
Give your agent the following goal:
/goal Optimize a local copy of
tirx-kernels/tirx_kernels/basic/fp16_bf16_gemm.pyfor FP16C = A @ B.T,M=N=K=1024, on NVIDIA B200. Aim to match or beattorch-cublaswhile preserving correctness. Compare original, optimized, and cuBLAS timings on the same GPU with identical settings over three independent runs. Report latencies, speedups, and any remaining gap.
Read correctness before speedup. Compare original, optimized, and cuBLAS timings on the same GPU with identical settings over three independent runs. Check the reported latencies, speedups, and any remaining gap alongside the modified kernel.
Run an existing workload¶
Follow Optimization Runs for one complete example: install the setup dependencies, select the FP16 GEMM workload, prepare a worktree, launch an agent, and inspect the resulting frontier.