{"product_id":"cuda-and-gpu-parallel-computing-engineering-accelerating-scientific-and-high-performance-workloads-through-cuda-kernels-memory-optimization-and-mul-9798196510748","title":"CUDA and GPU Parallel Computing Engineering: Accelerating Scientific and High-Performance Workloads Through CUDA Kernels, Memory Optimization, and Mul","description":"\u003cp\u003e • Author(s): Eamon Virek\u003cbr\u003e • Publisher: Independently Published\u003cbr\u003e • Publisher Imprint: Independently Published\u003cbr\u003e • BISAC: Programming - Parallel\u003c\/p\u003e\u003cp\u003e\u003c\/p\u003e\u003cp\u003e\u003cb\u003eA practical guide to high-performance CUDA development\u003c\/b\u003e for engineers, researchers, and developers who need more than introductory examples. This book focuses on the full workflow of GPU computing, from understanding how streaming multiprocessors execute warps to building maintainable, testable, and scalable applications for real scientific workloads.\u003c\/p\u003e\u003cp\u003eThe chapters move from core architecture and programming fundamentals into profiling, memory tuning, numerical accuracy, and multi-GPU scaling. You will see how to turn a correct kernel into an efficient one, how to measure bottlenecks with Nsight tools, and how to make informed tradeoffs between occupancy, bandwidth, latency, and precision.\u003c\/p\u003eWhat this book covers\u003col\u003e\n\u003cli\u003e\n\u003cb\u003eGPU architecture and execution behavior\u003c\/b\u003e, including warps, scheduling, memory hierarchy, and data movement costs.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eCUDA kernel design\u003c\/b\u003e, with launch configuration, indexing, synchronization, debugging, and reusable interfaces.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003ePerformance engineering\u003c\/b\u003e, using profiling metrics and iterative optimization based on measured results.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eMemory optimization\u003c\/b\u003e, including coalescing, shared memory tiling, register pressure, cache behavior, and data layout.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eCommon scientific patterns\u003c\/b\u003e, such as stencils, reductions, scans, sparse formats, and batched linear algebra.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eNumerical correctness\u003c\/b\u003e, with floating point behavior, stable summation, boundary handling, and CPU validation.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eAdvanced coordination techniques\u003c\/b\u003e, such as warp and block level operations, streams, events, and asynchronous overlap.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eHost and multi-GPU engineering\u003c\/b\u003e, covering pinned memory, unified memory, partitioning strategies, NCCL, halo exchange, and scaling studies.\u003c\/li\u003e\n\u003c\/ol\u003eWhy it stands out\u003cul\u003e\n\u003cli\u003e\n\u003cb\u003eEngineering-first approach\u003c\/b\u003e, centered on real optimization decisions rather than isolated syntax.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eWorkflow oriented\u003c\/b\u003e, with profiling, testing, benchmarking, and regression tracking built into the discussion.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eUseful for scientific computing\u003c\/b\u003e, especially stencil solvers, sparse methods, reductions, and iterative pipelines.\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eBuilt for maintainability\u003c\/b\u003e, with guidance on project structure, code reuse, and repeatable validation.\u003c\/li\u003e\n\u003c\/ul\u003e\u003cp\u003e\u003ci\u003eIdeal for anyone who wants to write CUDA code that is not only correct, but also fast, traceable, and ready for production-scale workloads.\u003c\/i\u003e\u003c\/p\u003e","brand":"Independently Published","offers":[{"title":"Paperback","offer_id":47892784873623,"sku":"9798196510748","price":1571.0,"currency_code":"INR","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0666\/3471\/1191\/files\/9798196510748.webp?v=1781188939","url":"https:\/\/atlanticbooks.com\/products\/cuda-and-gpu-parallel-computing-engineering-accelerating-scientific-and-high-performance-workloads-through-cuda-kernels-memory-optimization-and-mul-9798196510748","provider":"Atlantic Books","version":"1.0","type":"link"}