Job Modeling for Power Forecasting and Analysis on the Astra Supercomputer.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Technology for determining whether an inter-process type message has been successfully sent from a first process to a second process running on a single computer with a single processor(s) set. A variable (for example, a bit value) is used to indicate whether the inter-process message has been communicated between the processes. A timer and a predetermined timeout threshold are used to determine if the inter-process message has been pending for too long without being successfully communicated.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Our technology leverages artificial intelligence (AI) to enhance the user experience in High Performance Computing (HPC) environments. By analyzing user behavior and providing personalized recommendations, our AI system helps HPC users optimize their workflows and improve productivity. Additionally, we offer an advanced image similarity search feature, which utilizes AI algorithms to identify and retrieve visually similar images, saving users valuable time and effort in their research and analysis.
The Village of Questa, New Mexico is aiming to become a regional clean energy hub with robust and diverse employment opportunities for the local community supported by the energy sector and by other businesses inspired or attracted by abundant clean energy, outdoor recreation, and cultural opportunities. A coalition of stakeholders in the Village of Questa, comprising the Village, Kit Carson Electric Cooperative (KCEC), Questa Economic Development Fund, and Chevron, is exploring options to develop hydrogen production facilities as an opportunity to create jobs, provide reliable clean energy, and utilize former mine resources. Questa is home to a molybdenum mine owned by Chevron that closed in 2014. Several residents in Questa and surrounding communities lost their jobs when the mine closed and transitioned from active operations into environmental remediation. Although remediation efforts have been ongoing since 2014 and are expected to continue for at least 16 more years, the number of jobs with Chevron is much smaller now than it was before the closure. Between available workforce, brownfield land, and water rights formerly supporting mine operations but now in a transition period, there are considerable local resources that could be directed toward clean energy generation. Questa's electricity supply is already 100% solar during daylight hours thanks to Kit Carson Electric Cooperative's (KCEC's) strategic decision-making and partnering over the last decade. Now, Questa, KCEC, and Chevron are exploring the potential costs and benefits of siting an electrolytic hydrogen production facility and additional solar photovoltaic (PV) capacity in Questa to further advance the region's clean energy economy. In this report, we estimated the potential economic impacts (i.e., jobs, value added, gross output, tax revenue) of constructing and operating a combined hydrogen (32 MW polymer electrolyte membrane electrolizer + 7.5 MW fuel cell) and solar facility (22.5 MW) in the Village of Questa, as well as the resulting economic spillovers to Taos County and the state of New Mexico. We employ an input-output model that leverages IMPLAN's economic data for the region complemented by construction and operating expenses estimated by NREL and feedback from the local coalition to evaluate the direct, indirect and induced effects of the project construction (transient impacts) and operation (more permanent impacts). Based on the area's average trade profile, feedback from the coalition and current market conditions, these projects are expected to support 487 full-time equivalent jobs during construction, generating $\$24$ million in income for those workers and $\$82$ million in local economic activity in the state. Of those jobs, 106 are expected to be construction sector jobs. These investments are also estimated to add $\$36.5$ million to New Mexico's gross state product (GSP). In the Village of Questa, we estimate 16 jobs will be supported in construction and transportation industries, generating $\$0.9$ million in earnings. In Taos County, the construction phase is expected to support 285 jobs primarily in construction and professional services, while manufacturing jobs dominate the results for the Rest of New Mexico. The Village is also estimated to receive $\$0.9$ million in tax revenue from the construction phase alone. Once in operation, the project continues to impact the state and Questa. Around 20 jobs (full-time equivalent for each year of operation) are supported across New Mexico, with approximately 11 directly employed in Questa by both facilities. The total annual local economic activity supported by ongoing operations is just over $\$1.3$ million/yr, generating $\$1.6$ million/yr in additional income in the state. Annual operations are estimated to add $\$2.1$ million to the state's GSP. The Village is expected to receive around $\$43,000$/yr in tax revenue. Impacts vary significantly depending on which businesses are supplying materials, equipment and services, and where construction workers reside. Choosing local suppliers will most benefit Questa and the New Mexico economy, adding up to 500 jobs during construction and 13 long-term jobs. Local and state governments may consider ways to incentivize local businesses in order to maximize economic benefits.
The Village of Questa, New Mexico is aiming to become a regional clean energy hub with robust and diverse employment opportunities for the local community supported by the energy sector and by other businesses inspired or attracted by abundant clean energy, outdoor recreation, and cultural opportunities. A coalition of stakeholders in the Village of Questa, comprising the Village, Kit Carson Electric Cooperative (KCEC), Questa Economic Development Fund, and Chevron, is exploring options to develop hydrogen production facilities as an opportunity to create jobs, provide reliable clean energy, and utilize former mine resources. Questa is home to a molybdenum mine owned by Chevron that closed in 2014. Several residents in Questa and surrounding communities lost their jobs when the mine closed and transitioned from active operations into environmental remediation. Although remediation efforts have been ongoing since 2014 and are expected to continue for at least 16 more years, the number of jobs with Chevron is much smaller now than it was before the closure. Between available workforce, brownfield land, and water rights formerly supporting mine operations but now in a transition period, there are considerable local resources that could be directed toward clean energy generation. Questa's electricity supply is already 100% solar during daylight hours thanks to Kit Carson Electric Cooperative's (KCEC's) strategic decision-making and partnering over the last decade. Now, Questa, KCEC, and Chevron are exploring the potential costs and benefits of siting an electrolytic hydrogen production facility and additional solar photovoltaic (PV) capacity in Questa to further advance the region's clean energy economy. In this report, we estimated the potential economic impacts (i.e., jobs, value added, gross output, tax revenue) of constructing and operating a combined hydrogen (32 MW polymer electrolyte membrane electrolizer + 7.5 MW fuel cell) and solar facility (22.5 MW) in the Village of Questa, as well as the resulting economic spillovers to Taos County and the state of New Mexico. We employ an input-output model that leverages IMPLAN's economic data for the region complemented by construction and operating expenses estimated by NREL and feedback from the local coalition to evaluate the direct, indirect and induced effects of the project construction (transient impacts) and operation (more permanent impacts). Based on the area's average trade profile, feedback from the coalition and current market conditions, these projects are expected to support 487 full-time equivalent jobs during construction, generating $\$24$ million in income for those workers and $\$82$ million in local economic activity in the state. Of those jobs, 106 are expected to be construction sector jobs. These investments are also estimated to add $\$36.5$ million to New Mexico's gross state product (GSP). In the Village of Questa, we estimate 16 jobs will be supported in construction and transportation industries, generating $\$0.9$ million in earnings. In Taos County, the construction phase is expected to support 285 jobs primarily in construction and professional services, while manufacturing jobs dominate the results for the Rest of New Mexico. The Village is also estimated to receive $\$0.9$ million in tax revenue from the construction phase alone. Once in operation, the project continues to impact the state and Questa. Around 20 jobs (full-time equivalent for each year of operation) are supported across New Mexico, with approximately 11 directly employed in Questa by both facilities. The total annual local economic activity supported by ongoing operations is just over $\$1.3$ million/yr, generating $\$1.6$ million/yr in additional income in the state. Annual operations are estimated to add $\$2.1$ million to the state's GSP. The Village is expected to receive around $\$43,000$/yr in tax revenue. Impacts vary significantly depending on which businesses are supplying materials, equipment and services, and where construction workers reside. Choosing local suppliers will most benefit Questa and the New Mexico economy, adding up to 500 jobs during construction and 13 long-term jobs. Local and state governments may consider ways to incentivize local businesses in order to maximize economic benefits.
High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.
COVID-19 pandemic has affected clean energy labor market. Using real-time job vacancy data, this study analyzes the impacts of the pandemic on the U.S. clean energy labor market in 2020, including biomass, energy efficiency (EE), electric vehicle (EV), power/microgrid, solar, and wind industries. This study identifies how COVID-health factors and public health interventions influence clean energy job availability during the early COVID pandemic. Overall, California had the most energy jobs and experienced a significant decrease in April 2020. EV and solar had the highest percentages of job vacancies during the pandemic in general. Still, lockdowns had the most severe influence on EE and wind jobs. Stay-at-home orders negatively affected clean energy job vacancies in biomass, EV, power/microgrid, and wind. Social-gathering restrictions, however, did not have much influence. Increased COVID tests at the state level had the strongest and most positive influence on clean energy job postings, indicating the importance of a state's ability to manage public health infrastructure or crisis issues. COVID hospitalizations negatively influenced the job vacancies in biomass and wind but did not affect the other four sectors; conversely, as COVID death numbers increased, the number of jobs in biomass, EV, power grid, solar, and wind decreased, but not in EE jobs.
Job scheduling at supercomputing facilities is important for achieving high utilization of these valuable resources while ensuring effective execution of jobs submitted by users. The jobs are scheduled according to their specified resource demands such as expected job completion times, and the available resources based on allocations. Jobs that overrun their allocated times are terminated, for example, after a grace-period. It is non-trivial and often very complex for users to accurately estimate the completion times of their jobs, and consequently they face a dilemma: underestimate the job time to have a higher priority and risk job termination due to overrun, or overestimate it to ensure its completion and risk its delayed execution. In this paper, we investigate whether providing grace-period can benefit facility performance by developing a game- theoretic model between a facility provider and multiple users for a simplified scheduling scenario based on job execution times. We present closed-form expressions for the provider’s and user’s best-response strategies to maximize their respective utility functions. We describe conditions under which offering a grace-period is advantageous to both facility provider and users by deriving the Nash equilibrium of the game.
Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.
The power & energy demands of HPC machines have grown significantly. Modern exascale HPC systems require tens of megawatts of combined power for computing resources and cooling facilities at full capacity. The current energy trend is not sustainable for future HPC systems, and there is a need to work toward the energy efficiency aspect of HPC performance. Energy awareness of the HPC applications at the job level is essential for running an efficient HPC system. This work aims to develop a pipeline to provide a production-level system-wide overview of the HPC workloads' power profile while handling evolving workloads exhibiting new power trends. We developed an open-set classification model for HPC jobs based on the properties of power profiles to continuously provide a system-wide holistic view of recently completed jobs. The pipeline helps continuously monitor the job-level power usage pattern of HPC and enables us to capture the new trends in applications' power behavior. We employed a comprehensive set of techniques to generate job-level data, custom-designed feature extraction methods to extract critical features from jobs' power profiles, clustering techniques powered by generative modeling, and open-set classification for identifying job profiles into known classes or an unknown set. With extensive evaluations, we demonstrate the effectiveness of each component in our pipeline. We provide an analysis of the resulting clusters that characterize the power profile landscape of the Summit supercomputer from more than 60K jobs executed in a year. The open-set classification classifies the known data sets into known classes with high accuracy and identifies unknown data noints with over 85% accuracy.
Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.
Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).
AQDrop is a job management system designed to streamline access to the Advanced Quantum Testbed (AQT) at NERSC (National Energy Research Scientific Computing Center). It serves as a centralized middleware layer between researchers and quantum processing hardware. Key Features: AQDrop provides a FastAPI-based server backed by PostgreSQL for job submission, queue management, and role-based access control (members, operators, and administrators). Users submit Qiskit circuits via JSON payloads, which are queued, dispatched to the QPU through the Qubic API, and returned as measurement counts. A Python client library and web dashboard round out the interface options. Primary Use: Researchers submit quantum circuit jobs from a laptop or login node; an operator client executes those jobs on the AQT's physical QPU and returns results — all coordinated through the central API. Advantages: Compared to ad-hoc or direct hardware access, AQDrop adds structured queue management, auditable job-status tracking and OAuth2 authentication — reducing scheduling conflicts and unauthorized access. Its containerized deployment also improves reproducibility and scalability. Overall, AQDrop functions as a purpose-built quantum job broker tailored to NERSC's specific hardware and institutional access requirements.
In the past decade, solar power, with and without energy storage has become the fastest growing source of energy generation in the world. In the U.S., solar employment more than doubled from 105,145 jobs in 2011 to 255,037 jobs in 2021, four times faster than the U.S. job growth rate overall. These factors, combined with technology advancements, creates a skills gap that puts tremendous stress on society to deliver the workers to fill the open job requisitions. The Cyberguardians and STEM Warriors project (Cyberguardians) was designed to address the trained-worker shortage in the energy industry in three ways: 1) by developing educational curriculum that addresses DER technology changes; 2) by delivering curriculum to prospective workers, including military veterans and their families, via universities, community colleges, and vocational training outlets; and 3) introducing individuals who have completed training to employers that can hire them. Cyberguardians exceeded its curriculum goals by producing 27 academic units of university-accredited material (12 total courses) covering energy fundamentals, smart inverters, Distributed Energy Resource (DER) data communication, cybersecurity, standardization, certification, data analytics, and IEEE 1547 standard topics. The North American Board of Certified Energy Practitioners (NABCEP) also accredited the material for use in their credential program. Seven instructors were recruited and trained, and six academic institutions (University of California San Diego, State University of New York, North Carolina State University, Harper Community College, Green Village Academy, and the SunSpec Alliance) were enlisted, meeting program goals. All course material was published under the Creative Commons license and made available royalty free, thus providing a long-lasting public benefit. The program’s outreach program vastly exceeded program goals and incorporated the efforts of 13 outreach partners (11 of which are veteran focused), an advisory board representing 15 companies, webinars and 10’s of thousands of email messages sent to prospective students and hiring managers. Despite these efforts, the global pandemic depressed anticipated program participation by about a third. Still, a total of 396 students enrolled and 289 completed the courses and were accredited. The job applicant task achieved similar results (111 realized vs a 174 goal) but reported job placement was weaker at (9 realized vs. a 51 goal). The Cyberguardians program fills a critical void for cost-effective, royalty-free curriculum and training pertaining to DER technologies and cybersecurity that prospective energy workers must possess to be effective in the 21 st century. On this basis alone, the investment of taxpayer funds will pay dividends for years to come.
Fermilab is the first High Energy Physics institution to transition from X.509 user certificates to authentication tokens in production systems. All the experiments that Fermilab hosts are now using JSON Web Token (JWT) access tokens in their grid jobs. Many software components have been either updated or created for this transition, and most of the software is available to others as open source. The tokens are defined using the WLCG Common JWT Profile. Token attributes for all the tokens are stored in the Fermilab FERRY system which generates the configuration for the CILogon token issuer. High security-value refresh tokens are stored in Hashicorp Vault configured by htvault-config, and JWT access tokens are requested by the htgettoken client through its integration with HTCondor. The Fermilab job submission system jobsub was redesigned to be a lightweight wrapper around HTCondor. The grid workload management system GlideinWMS which is also based on HTCondor was updated to use tokens for pilot job submission. For automated job submissions a managed tokens service was created to reduce duplication of effort and knowledge of how to securely keep tokens active. The existing Fermilab file transfer tool ifdh was updated to work seamlessly with tokens, as well as the Fermilab POMS (Production Operations Management System) which is used to manage automatic job submission and the RCDS (Rapid Code Distribution System) which is used to distribute analysis code via the CernVM FileSystem. The dCache storage system was reconfigured to accept tokens for authentication in place of X.509 proxy certificates. As some services and sites have not yet implemented token support, proxy certificates are still sent with jobs for backwards compatibility, but some experiments are beginning to transition to stop using them.