Skip to content

Booksellers & Trade Customers: Sign up for online bulk buying at trade.atlanticbooks.com for wholesale discounts

Booksellers: Create Account on our B2B Portal for wholesale discounts

GPU Programming using Rust and CUDA: Exploring Rust's potential in GPU and parallel computing using Rust-CUDA, cuda-oxide, and RustaCUDA

by Maris Fenlor
Sold out
₹8,623.00
Original price ₹8,623.00
Original price ₹8,623.00
₹8,623.00
Current price ₹8,623.00

Imported Edition - Ships in 18-21 Days

Free Shipping in India on orders above Rs. 500

Request Bulk Quantity Quote
+91
Book cover type: Paperback
  • ISBN13: 9789349174375
  • Binding: Paperback
  • Subject: N/A
  • Publisher: Gitforgits
  • Publisher Imprint: Gitforgits
  • Publication Date:
  • Pages: 166
  • Original Price: USD 79.99
  • Language: English
  • Edition: N/A
  • Item Weight: 295 grams
  • BISAC Subject(s): Programming / Algorithms

C++ has been the go-to for GPU programming for almost 20 years. Can Rust do the job, and how well?

This book is all about getting hands-on with different toolchains that connect Rust to NVIDIA hardware. There's RustaCUDA for safe host-side control, the Rust-CUDA project for writing kernels in pure Rust, and NVIDIA's experimental cuda-oxide compiler with its typed launches and async execution graphs.

We're going to build one Cargo workspace that keeps on growing. It'll include device queries, launch planning, Rust-written kernels, memory optimization, parallel reductions and scans, multi-stream pipelines, matrix multiplication benchmarked against cuBLAS, a Monte Carlo option pricer validated against a closed formula, and a complete batched inference application measured against a Python baseline. We'll check every result against a CPU reference, and the reports will give accurate numbers, including where libraries outperform hand-written kernels and where experimental toolchains are still a work in progress.

Key Learnings

Launch, synchronize, and verify GPU kernels with ownership-managed device memory.

Write real CUDA kernels using Rust-CUDA and cuda-oxide.

Plan grids, blocks, and warps for 2D workloads.

Accelerate transfer speeds with pinned memory and coalesced access patterns.

Build race-free thread cooperation using shared memory, barriers, and atomics.

Overlap transfers with computation using streams, events, and async Rust pipelines.

Optimize matrix multiplication and benchmark against cuBLAS ceiling.

Wrap CUDA C library safely with handles, error enums, and Drop.

Ship complete batched GPU inference application against Python baselines.

Diagnose performance with Nsight Systems, Nsight Compute, and compute-sanitizer.

Table of Content

New Beneficiary of GPU Computing

Thinking in Threads

Commanding GPU

Writing GPU Kernels

Cleaner Kernels with cuda-oxide

Mastering GPU Memory

Making Threads Cooperate

Keeping GPU Busy

Delivering Real Math

Borrowing NVIDIA's Muscle

Shipping Complete GPU Application

Proving Performance

Trusted for over 49 years

Family Owned Company

Secure Payment

All Major Credit Cards/Debit Cards/UPI & More Accepted

New & Authentic Products

India's Largest Distributor

Need Support?

Whatsapp Us