
Resources for Professionals Interested in High Performance Computing
For aspiring and early career HPC professionals, this rapidly evolving HPC/AI environment presents considerable challenges, including demands for specialized AI infrastructure skills, increasingly complex architectures and libraries, and a shift toward lower-precision computing
Page Content:
Once used almost exclusively by governments and researchers, high-performance computing (HPC) is now converging with AI and offering enticing possibilities for users across industry and commercial sectors.
At Supercomputing 2025’s “Future of HPC” panel, moderator Wim Slagter recalls how Todd Simons, an HPC expert at Rolls-Royce Americas, described exascale computing access as being akin to accessing “a time machine … you’re jumping into the future of where computers will be years from now and getting your applications ready.”
For aspiring and early career HPC professionals, this rapidly evolving HPC/AI environment presents considerable challenges, including
- Demands for specialized AI infrastructure skills
- Increasingly complex architectures and libraries
- A shift toward lower-precision computing
Even in tumultuous times, however, knowledge remains an essential power.
Here, you’ll learn…
- Which HPC skills are foundational today? Master fundamentals such as parallel and distributed computing, AI–simulation workflow integration, and performance optimization, with an eye toward next-gen system designs.
- What trends are shaping HPC’s future? CPU-based clusters are giving way to hybrid heterogeneous architectures, and the classical vs. quantum mindset is shifting to classical + quantum possibilities.
- What challenges is the HPC field facing? Urgent issues include AI integration, power consumption, and system scalability and complexity.
- What are promising career paths in HPC? HPC skills are netting salary boosts in many sectors; in-demand roles include cluster engineers, parallel computing engineers, and HPC–AI systems architect.
- Which ethical challenges are most urgent? Topping the list are energy consumption, equitable HPC access, and accountability.
- How can I stay up-to-date on HPC news and research? Access the latest standards, SME insights, and industry trends.
Learn more about the field at The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC)
HPC: The Fundamentals
The shift toward hyperscale AI platforms is reshaping the HPC landscape across hardware, software, and economics. As Daniel Reed and colleagues put it in their article on “Scientific Computing in an AI World,” HPC’s challenge is to “ride the wave” of AI-driven infrastructure while also building the future, including through ambitious, next-generation system designs that
- Use energy-aware algorithms to reduce energy per validated scientific outcome
- Innovate architectures focused on memory and interconnect efficiency
- Build software stacks optimized for hybrid AI/simulation workflows
HPC Skills to Focus on Now
- Parallel and distributed computing: Designing software that efficiently scales across CPUs, GPUs, clusters, and heterogeneous systems.
- AI/simulation workflow integration: Bridging machine learning and traditional simulation.
- Performance optimization and profiling: Including identifying bottlenecks in memory, storage, and networking to improve efficiency at scale.
- Energy- and resource-aware computing: Optimizing for both performance and energy efficiency.
What Are Key Trends in HPC?
As Matt Vincent emphasizes in the Data Center Frontier report, Present and Future HPC Trends, Forecast, and Implications Within Data Centers, AI and HPC workloads are increasingly hosted side by side. This is one of several developments creating new opportunities—and challenges—in HPC today.
Heterogeneous and AI-Accelerated Architectures
HPC systems are moving from CPU-based clusters to hybrid heterogeneous architectures combining CPUs, GPUs, and specialized AI processors. The goal is to maximize performance and power efficiency for AI-driven workloads for tasks including
- scientific simulation
- AI training/inference
- data analytics
- real-time modeling
What Do Hybrid Heterogeneous Architectures Mean for HPC?
- Increased convergence between HPC and AI infrastructure
- Increased use of AI-assisted simulation and modeling
- Demand for software that can run across multiple processor types
- Increased commercial relevance of HPC beyond academia and government labs
What Are the Technical Challenges of Heterogeneous Architectures?
- Programming across heterogeneous hardware
- Portability issues between vendors and architectures
- Memory bandwidth bottlenecks
- Efficient scheduling of mixed CPU/GPU workloads
- Energy and cooling requirements at exascale scale
- Maintaining numerical precision in AI/HPC hybrid workflows
Where Can I Learn More?
Hybrid Quantum–HPC
In the hybrid quantum-HPC paradigm, classical supercomputers work closely with quantum processors (QPUs) and treat the quantum component as an accelerator for complex tasks.
These hybrid architectures combine
- CPUs for orchestration and control
- GPUs for acceleration and error correction
- QPUs for highly specialized computational tasks
As “Towards A Unified Quantum Platform” notes, 71% of HPC centers globally plan to deploy quantum computing in some form in 2026, shifting away from the classical vs. quantum model to emerging classical + quantum possibilities.
What Does Hybrid Quantum–HPC Mean for HPC?
In addition to the growth of quantum-centric supercomputing, this trend is likely to result in
- Workflows combining simulation, AI, and quantum optimization
- Development of quantum-aware schedulers and middleware
- Increased investment in HPC infrastructure that supports quantum systems
Where Can I Learn More?
- Collaboration Accelerates U.S. Leadership in Hybrid Quantum–Classical Computing
- Shifting Sands of Hardware and Software in Exascale Quantum Mechanical Simulations
- Architecting a Full-Stack Superconducting Fault-Tolerant Quantum Computer
- IBM Releases a New Blueprint for Quantum-Centric Supercomputing
Exascale Computing and the Energy/Data Bottleneck
HPC systems are entering an exascale era where performing at a quintillion calculations per second (or more) may soon be the norm. Such raw compute power creates considerable challenges, not least in
- Energy consumption
- Storage throughput
- Data movement
- Interconnect efficiency
As systems scale upward, moving and storing data often costs more than computation itself.
What Does Exascale Computing Mean for HPC?
Among the implications, exascale computing results in
- Greater emphasis on energy-efficient architectures
- Expansion of liquid cooling and advanced interconnect technologies
- Increased focus on storage hierarchy and memory optimization
- More regional/national investment in HPC infrastructure
Technical Challenges
- Massive power consumption
- Thermal management and cooling complexity
- Storage bandwidth bottlenecks
- Fault tolerance across millions of parallel processing elements
- Efficient data movement between compute nodes
- Reliability and resilience in very large distributed systems
Where Can I Learn More?
What Are the Challenges in HPC Today?
As is true across the tech spectrum, the increasing intertwining of AI in HPC entails multiple challenges, including the following.
Energy Consumption and Cooling
Power delivery, cooling, and carbon footprint have become limiting factors on HPC, especially exascale and AI-focused systems, which often require tens of megawatts to operate efficiently. Research is therefore focusing on improvements in power management, liquid cooling, and energy-aware scheduling to keep systems sustainable.
Where Can I Learn More?
Scalability and System Complexity
HPC increasingly combines CPUs, GPUs, accelerators, huge memory pools, and fast interconnects across thousands of nodes. This makes efficiently scaling applications difficult due to communication delays, memory bottlenecks, and hardware failures. All of these issues escalate as systems grow larger. The complexity of these exascale systems leaves software and programming models struggling to keep up.
Where Can I Learn More?
Integrating AI Workloads With Traditional HPC
AI and machine-learning workloads are changing how HPC systems are designed and used. While traditional scientific simulations require high numerical precision, AI workloads prioritize massive parallelism and accelerator performance. Combining both goals efficiently on shared infrastructure creates challenges in many areas, including scheduling, storage, networking, software frameworks, and resource management.
Where Can I Learn More?
Building a Career in HPC Today
HPC’s rapid convergence with AI is making GPU systems engineering, distributed computing, storage architecture, and performance optimization among the fastest-growing technical specialties in computing today.
In addition to growth in traditional HPC markets such as government and defense, HPC adoption in the life science sector is expected to grow dramatically over the next few years to accelerate drug discovery for various diseases.
According to Fortune Business Insights, the global HPC market was valued at US$59.33 billion in 2025 and is projected to grow to US$128.84 billion by 2034. This growth is fueled by several factors, including the increasing number of data centers and cloud services, and adoption of quantum computing and cloud-based HPC.
All of these changes bode well for the HPC job market. Following are some key roles to consider.
HPC Systems Administrator/Cluster Engineer
- Focus: Managing Linux-based HPC clusters, schedulers, storage, and user environments
- Top sectors: National labs, universities, biotech, aerospace, and semiconductor design
- Job titles: HPC Systems Administrator, Cluster Engineer, Research Computing Engineer
- Requirements: Linux administration, scripting, and networking fundamentals; Slurm/PBS familiarity; Linux+ and Python/Bash skills
- Key certifications: Software Professional Certification—Level 1, Professional Software Engineering Master Certification, Red Hat system administrator/engineer certifications (RHCSA/ RHCE)
- Average US salary: USD$88,927 (ZipRecruiter)
- Where to Network: The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), IEEE International Conference on Cluster Computing
- Related research: IEEE Transactions on Parallel and Distributed Systems, IEEE Internet Computing
HPC Software Developer/Parallel Computing Engineer
- Focus: Developing scalable scientific and AI applications using parallel programming
- Top sectors: Scientific computing, climate modeling, defense, pharmaceuticals, AI infrastructure
- Job titles: HPC Software Engineer, Parallel Computing Developer, Scientific Programmer
- Requirements: Strong CS/math background; GPU programming experience; and experience with C/C++, MPI, OpenMP, CUDA, and debugging/profiling tools
- Key certifications: Professional Software Engineering Master Certification, Fundamentals of Accelerated Computing with Modern CUDA C++, NVIDIA Deep Learning Institute (DLI), HPC Certification Forum
- Average US salary: USD$147,524 (ZipRecruiter)
- Where to Network: The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), IEEE International Parallel & Distributed Processing Symposium
- Related research: IEEE Transactions on Parallel and Distributed Systems, IEEE Software
AI Infrastructure/GPU Platform Engineer
- Focus: Building and optimizing GPU clusters for AI and HPC convergence
- Top sectors: AI companies, cloud providers, autonomous systems, finance, hyperscalers
- Job titles: AI Infrastructure Engineer, GPU Platform Engineer, ML Systems Engineer
- Requirements: Distributed system knowledge; expertise with NVIDIA CUDA; and experience with Kubernetes, GPUs, RDMA/InfiniBand, and cloud/HPC integration
- Key credentials: Certified Kubernetes Administrator (CKA), Professional Software Engineering Master Certification, NVIDIA-Certified Professional AI Infrastructure
- Average US salary: USD$127,066 (ZipRecruiter)
- Where to Network: The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), IEEE International Conference on High Performance Computing and Communications (HPCC)
- Related research: IEEE Transactions on Parallel and Distributed Systems, IEEE Transactions on Cloud Computing
HPC Data/Storage Engineer
- Focus: Managing high-throughput storage systems and large-scale scientific data movement
- Top sectors: Genomics, energy, weather modeling, AI training infrastructure
- Job titles: HPC Storage Engineer, Parallel File Systems Engineer, Data Infrastructure Engineer
- Requirements: Linux expertise; distributed systems experience; and experience with parallel storage systems, Lustre/GPFS, networking, and distributed storage architecture
- Key certifications: Professional Software Engineering Master Certification, Fundamentals of Accelerated Computing with Modern CUDA C++, NVIDIA Deep Learning Institute (DLI), HPC Certification Forum
- Average US salary: USD$116,916 (ZipRecruiter)
- Where to Network: The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), IEEE International Conference on High Performance Computing and Communications (HPCC)
- Related research: IEEE Transactions on Big Data, IEEE Transactions on Parallel and Distributed Systems
HPC–AI Systems Architect
- Focus: Designing next-generation hybrid AI/HPC infrastructure and large-scale compute strategy
- Top sectors: Hyperscalers, cloud providers, defense, national supercomputing centers
- Job titles: HPC Architect, AI Infrastructure Architect, Chief HPC Engineer
- Requirements: Experience with large-scale systems design, networking, GPUs, and cloud/HPC integration; leadership experience; advanced systems architecture experience; and HPC domain expertise
- Key certifications: Professional Software Engineering Master Certification, NVIDIA Deep Learning Institute (DLI), HPC Certification Forum, CompTIA Cloud+
- Average U.S. salary range: USD$128,756 (ZipRecruiter)
- Where to Network: The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), IEEE International Symposium on High-Performance Computer Architecture (HPCA)
- Related research: IEEE Transactions on Parallel and Distributed Systems and IEEE Transactions on Cloud Computing
Ethical Issues in HPC
- In its 2025 Energy and AI report, the International Energy Agency notes that “without energy, there is no AI.” While AI might well be able to eventually help us to transform energy use in beneficial ways, in the near term, it simply needs more. Much more. As IEA points out, while a typical data center consumes as much electricity as 100,000 households, the next-generation campuses being built to support AI’s growth will demand 20 times that amount.Energy is just one ethical issue HPC must face, but it is a massive one that intertwines with many others.
Energy Consumption of HPC–AI Infrastructure
“Compute efficiency” has grown from a merely technical concern to an ethical one. Training frontier AI models increasingly requires hyperscale GPU clusters and large data centers; without sustainable energy strategies, those data centers will strain regional power grids and increase carbon emissions in ways that are currently difficult to calculate.
Whether unlimited scaling of computational workloads is environmentally sustainable is just one of the questions that researchers, governments, and other stakeholders are grappling with here.
Other key ethical questions:
- Who will pay the environmental costs of hyperscale computing?
- Should organizations be required to report compute-related emissions and water use?
- Should sustainability issues put limits on AI/HPC growth?
Where Can I Learn More?
Concentration of Computing Power
A relatively small group of hyperscalers, governments, universities, and major tech firms control most advanced HPC and AI infrastructure. This strategic advantage impacts scientific research, economic competitiveness, defense capabilities, and AI development itself.
This increasing concentration of compute results in “compute inequality” as smaller institutions, developing nations, independent researchers, and nonprofits—and the range of ideas they offer and interests they serve—are excluded from meaningful participation in frontier research.
Key ethical questions:
- Who gets access to advanced computational infrastructure?
- Does concentration of compute reduce scientific openness and innovation?
- Could compute access become a geopolitical or economic gatekeeping mechanism?
Where Can I Learn More?
Responsible Use of HPC
Among the things that advanced surveillance systems, autonomous weapons research, cyber operations, large-scale biometric analysis, and AI-driven decision-making systems have in common? They are all increasingly controlled by HPC systems.
HPC infrastructure enables increasingly autonomous AI models to operate at unprecedented scale and speed. Ethically, it’s no longer a question of whether these systems can be built; they are here and the issues only begin with accountability to governance.
Key Ethical Questions
- What limits should exist on military and surveillance applications of HPC?
- How should accountability function in autonomous systems?
- Can democratic oversight keep pace with large-scale computational capability?
Where Can I Learn More?
Resources: The HPC Knowledge Hub
Stay up-to-date with the latest on learning technologies by accessing our Tech News blog, which is updated daily with the insights, trends, and research related to all things computing. Among recent related articles are the following:
- Shaping the Future of HPC through Architectural Innovation and Industry Collaboration
- Who Benefits from AI? The State of AI Governance
- Shaping the Future of HPC through Architectural Innovation and Industry Collaboration
- Parallel Systems, Leadership, and Research Strategy in Computing: An Interview with Jean-Luc Gaudiot
- Exploring the Differences Between Parallel and Distributed Computing