CUDA ecosystem wrappers for Rust
// readme
baracuda
About the name. Yes, we know — it’s spelled barracuda (two Rs). That name was taken on crates.io, so we dropped one R and kept swimming.
A unified Rust ML-op facade over the NVIDIA CUDA ecosystem.
What baracuda is
baracuda is a Rust workspace that exposes every primitive an ML framework
expects — the union of PyTorch (torch.* + nn.functional) and JAX
(jax.lax.* + jax.numpy.*) — through a single Plan-based crate surface
called [baracuda-kernels]. Internally each plan dispatches to:
- The appropriate NVIDIA-library wrapper crate (cuBLAS, cuDNN, cuFFT, cuSOLVER, cuRAND, cuSPARSE, cuTENSOR, NPP, CV-CUDA, CUTLASS) when one already covers the op well, or
- A bespoke hand-rolled
.cukernel shipped in [baracuda-kernels-sys] when no NVIDIA library covers the op (or covers it poorly at the shapes that matter for modern transformer / vision / GNN workloads).
Callers import one crate (baracuda-kernels) and reach for one API
style. The dispatch decision — which is observable through
Plan::sku() for telemetry — is otherwise invisible. Switching from a
CUTLASS-backed SKU to a bespoke-backed SKU is a layout flag, not an…