Book description
Problem-Solving in High Performance Computing: A Situational Awareness Approach with Linux focuses on understanding giant computing grids as cohesive systems. Unlike other titles on general problem-solving or system administration, this book offers a cohesive approach to complex, layered environments, highlighting the difference between standalone system troubleshooting and complex problem-solving in large, mission critical environments, and addressing the pitfalls of information overload, micro, and macro symptoms, also including methods for managing problems in large computing ecosystems.
The authors offer perspective gained from years of developing Intel-based systems that lead the industry in the number of hosts, software tools, and licenses used in chip design. The book offers unique, real-life examples that emphasize the magnitude and operational complexity of high performance computer systems.
- Provides insider perspectives on challenges in high performance environments with thousands of servers, millions of cores, distributed data centers, and petabytes of shared data
- Covers analysis, troubleshooting, and system optimization, from initial diagnostics to deep dives into kernel crash dumps
- Presents macro principles that appeal to a wide range of users and various real-life, complex problems
- Includes examples from 24/7 mission-critical environments with specific HPC operational constraints
Table of contents
- Cover
- Title page
- Table of Contents
- Copyright
- Dedication
- Preface
- Acknowledgments
- Introduction: data center and high-end computing
- Chapter 1: Do you have a problem?
- Chapter 2: The investigation begins
- Chapter 3: Basic investigation
- Chapter 4: A deeper look into the system
- Chapter 5: Getting geeky – tracing and debugging applications
- Chapter 6: Getting very geeky – application and kernel cores, kernel debugger
- Chapter 7: Problem solution
- Chapter 8: Monitoring and prevention
- Chapter 9: Make your environment safer, more robust
- Chapter 10: Fine-tuning the system performance
- Chapter 11: Piecing it all together
- Subject Index
Product information
- Title: Problem-solving in High Performance Computing
- Author(s):
- Release date: September 2015
- Publisher(s): Morgan Kaufmann
- ISBN: 9780128010648
You might also like
book
Contemporary High Performance Computing
Contemporary High Performance Computing: From Petascale toward Exascale focuses on the ecosystems surrounding the world’s leading …
book
PacketCable Implementation
PacketCable Implementation Design, provision, configure, manage, and secure tomorrow's high-value PacketCable networks Jeff Riddel, CCIE® No. …
book
Optimizing Linux® Performance: A Hands-On Guide to Linux® Performance Tools
The first comprehensive, expert guide for end-to-end Linux application optimization Learn to choose the right tools—and …
book
Computer Games and Software Engineering
Computer games represent a significant software application domain for innovative research in software engineering techniques and …