A high-throughput and memory-efficient inference and serving engine for LLMs
Cuda
Verified open-source repositories and technical documentation formatted for LLM context windows and autonomous AI coding agents (42 repositories indexed).
The open-source AI voice studio. Clone, dictate, create.
SGLang is a high-performance serving framework for large language models and multimodal models.
World's fastest and most advanced password recovery utility
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.
CUDA on non-NVIDIA GPUs
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Open3D: A Modern Library for 3D Data Processing
NumPy & SciPy for GPU
A General-purpose Task-parallel Programming System in C++
πLeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginnersπ, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.π
The most powerful local music generation model that outperforms almost all commercial alternatives, supporting Mac, AMD, Intel, and CUDA devices.
NumPy aware dynamic Python compiler using LLVM
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
CUDA Templates and Python DSLs for High-Performance Linear Algebra
The Open-Source Elevenlabs alternative AI Voice Clone, Dub, Dictate, Transcribe, Audiobook creator and Voice workflow studio.
cuDF - GPU DataFrame Library
Containers for machine learning
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
A fast, scalable, high performance Gradient Boosting on Decision Trees library, used for ranking, classification, regression and other machine learning tasks for Python, R, Java, C++. Supports computation on CPU and GPU.
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Go package for computer vision using OpenCV 4 and beyond. Includes support for DNN, CUDA, OpenCV Contrib, and OpenVINO.
An interactive NVIDIA-GPU process viewer and beyond, the one-stop solution for GPU process management.
A Python framework for GPU-accelerated simulation, robotics, and machine learning.
FlashInfer: Kernel Library for LLM Serving
Making it easier to work with shaders
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
ALIEN is a CUDA-powered artificial life simulation program.
LWJGL is a Java library that enables cross-platform access to popular native APIs useful in the development of graphics (OpenGL, Vulkan, bgfx), audio (OpenAL, Opus), parallel computing (OpenCL, CUDA) and XR (OpenVR, LibOVR, OpenXR) applications.
cuML - RAPIDS Machine Learning Library
Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.
CUDA accelerated rasterization of gaussian splatting
A PyTorch Library for Accelerating 3D Deep Learning Research
ArrayFire: a general purpose GPU library.
AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
Optimized primitives for collective multi-GPU communication
Fast inference engine for Transformer models
Lightning fast C++/CUDA neural network framework
HIP: C++ Heterogeneous-Compute Interface for Portability
A deep learning package for many-body potential energy representation and molecular dynamics
CUDA Accelerated Robot Library