Change the repository type filter
All
Repositories list
67 repositories
- Diffusion model inference benchmarking, profiling and optimizing
TraceLens
PublicAutomating analysis from trace filesGEAK
Public- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs
Primus-SaFE
PublicPrimus-SaFE(Stability and Fault Endurance)- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs
- More token goodput from frontier models. A distributed, SLA-aware serving mesh — disaggregated prefill/decode, KV-aware routing, and cache offload, tuned to you…
Instella-MoE
Public- A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
Apex
PublicAgents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, o…ALTO
PublicALTO: Advanced Low-precision Training and OptimizationAgentKernelArena
PublicAgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, …pr_pundit
PublicAn agent for OSS contributionsvllm-2026
Publicmlperf-common
PublicFarSkip-Collective
PublicTraining and inference implementation of FarSkip-Collective models enabling communication-computation overlapmaxtext-slurm
PublicHummingbirdXT
PublicThis repository presents an efficient acceleration pipeline for Diffusion Transformer (DiT) based video generation models, optimized for AMD client-grade GPUs, …torchtitan-amd
PublicA PyTorch native platform for training generative AI modelsgpt-fast
PublicInstella-Math
PublicAMD-LLM
PublicTraining code and resources for AMD-135M language models on AMD GPUs.m3d_rocm
PublicThis project is an optimized version of Matrix3D. It has better compatibility with ROCm ecosystem.prime_amd
PublicPARD
PublicPARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)ReasonLite
PublicEfficient Reasoning Modelssd
PublicA lightweight inference engine supporting speculative speculative decoding (SSD).DynamicChunkingDiT
PublicOfficial code for DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunkingawesome-rocm-autodrive
Public
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.