Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Jobs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

MARBLE: A Multi-GPU Aware Job Scheduler for Deep Learning on HPC Systems

Deep learning (DL) has become a key tool for solving complex scientific problems. However, managing the multi-dimensional large-scale data associated with DL, especially atop extant multiple graphics processing units (GPUs) in modern supercomputers poses significant challenges. Moreover, the latest high-performance computing (HPC) architectures bring different performance trends in training throughput compared to the existing studies. Existing DL optimizations such as larger batch size and GPU locality-aware scheduling have little effect on improving DL training throughput performance due to fast CPU-to-GPU connections. Additionally, DL training on multiple GPUs scales sublinearly. Thus, simply adding more GPUs to a system is ineffective. To this end, we design MARBLE, a first-of-its-kind job scheduler, which considers the non-linear scalability of GPUs at the intra-node level to schedule an appropriate number of GPUs per node for a job. By sharing the GPU resources on a node with multiple DL jobs, MARBLE avoids low GPU utilization in current multi-GPU DL training on HPC systems. Our comprehensive evaluation in the Summit supercomputer shows that MARBLE is able to improve DL training performance by up to 48.3% compared to the popular Platform Load Sharing Facility (LSF) scheduler. Compared to the state-of-the-art of DL scheduler, Optimus, MARBLE reduces the job completion time by up to 47%.

Han, Jingoo↗

Understanding the Interplay between Hardware Errors and User Job Characteristics on the Titan Supercomputer

Designing dependable supercomputers begins with an understanding of errors in real-world, large-scale systems. The Titan supercomputer at Oak Ridge National Laboratory provides a unique opportunity to investigate errors when an actual system is actively used by multiple concurrent users and workloads from diverse domains at varying scales. This study presents a thorough analysis of 6, 908, 497 hardware errors from 18, 688 compute nodes of Titan for 312, 215 user jobs over a 3-year time period. Through careful joining of two system logs – the Machine Check Architecture (MCA) log and the job scheduler log – we show the correlated pattern of hardware errors for each job and user, in addition to individual descriptive statistics of errors, jobs, and users. Since the majority of hardware errors are memory errors, this study also shows the importance of error correcting in memory systems.

Lim, Seung-Hwan↗

Using Pilot Jobs and CernVM File System for Simplified Use of Containers and Software Distribution

High Energy Physics (HEP) experiments entail an abundance of computing resources, i.e. sites, to run simulations and analyses by processing data. This requirement is fulfilled by local batch farms, grid sites, private/commercial clouds, and supercomputing centers via High Throughput Computing (HTC). The growing needs of such experiments and resources being prone to trends of heterogeneity make it difficult for physicists to handle these resources directly. Additionally, HEP collaborations heavily rely on data and software releases, typically in the order of tens of gigabytes, while conducting simulations and analyses. Hence, aspects of scalability, reliability, and maintenance become crucial with regards to the distribution of the necessary data and software stack. The GlideinWMS [4] framework helps with the resource management problem by using pilot jobs, aka Glideins, to provision reliable elastic virtual clusters. Glideins are submitted to unreliable heterogeneous resources which are validated and customized by the Glideins to make the worker nodes available for end-user job execution. On the other hand, the CernVM File System (CernVM-FS or CVMFS) [1] helps with data distribution. It is a write-once, read-everywhere filesystem used to deploy scientific software to thousands of nodes on a worldwide distributed computing infrastructure. CVMFS is based on the Hyper Text Transfer Protocol and has been widely used within the particle physics community for (1) distributing experiment software and data such as calibrations, and (2) facilitating containerization by efficiently hosting container images along with providing containerization software, especially Singularity [3] GlideinWMS relies on CVMFS installed locally on the computing resources to satisfy the experiments' software needs. This requires system administrators' effort to install and maintain CVMFS at the sites and limits the use of sites, especially HPC resources, that do not have CVMFS installed. This poster presents a solution, taking advantage of Glideins to provide CVMFS at most sites without the need for a local installation. Doing so expands the pool of resources available for HEP experiments and reduces the effort of system administrators for current resources. Additionally, the proposed solution allows GlideinWMS to also start Singularity [3], a containerization software that can run unprivileged, on sites where neither CVMFS nor Singularity are available, including HPC sites. The benefits provided by this solution are: (1) lower overhead for site administrators in that they have less software to install, (2) an expanded pool of resources that run user jobs with easy access to software and data provided by CVMFS, thus making life easier for the scientists, and (3) improved flexibility to use HPC resources by enabling GlideinWMS pilot jobs to support HPC sites.

Urs, Namratha↗

Flexible Pilot Jobs Framework for Distributed High Throughput Computing

Experimental particle physics has been at the forefront of analyzing the world’s largest datasets for decades. The high-energy physics (HEP) community was among the first to develop suitable software and computing tools for this purpose. GlideinWMS is a Glidein-based workload management system whose purpose is to provide experiments like CMS at CERN, DUNE at Fermilab, and others, a way to access and efficiently use vast amounts of computing resources. This system wants to provide a simple way to submit jobs to a set of computing resources, that will be provided to users behind the scenes. Glideins are the pilot jobs executed on the worker nodes at the grid sites, performing operations such as hardware detection, environment setup, and error handling. After all these operations, they will launch the actual user job. Many grid sites are supported, such as shared clusters, Google CE, and AWS. My internship aimed to design and code a flexible pilot jobs framework that will replace the one used by GlideinWMS, developing a modular and flexible skeleton of the Glidein and adding further functionalities. My project also focused on the application of machine learning techniques as support to this management system.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ohio's Clean Energy Jobs Potential Through 2030

According to the U.S. Census Bureau, Ohio had 7,516,303 people in its working population (15 to 64 years of age) in 2019. The graphs below show solar photovoltaic (PV), land-based wind, battery energy storage (BES), and energy efficiency job estimates in 2020, 2025, and 2030. These job estimates do not represent net job creation. Rather, they represent the size of the workforce required to achieve projected national deployment levels of each technology for 2025 and 2030 if the state capture

clean energy↗

Regulators’ Energy Transition Primer: Economic Impacts of the Energy Transition on Energy Communities, Environmental Justice Considerations, and Implications on Clean Energy Jobs

Applications of new technology, such as horizontal drilling and hydraulic fracturing, enabled the United States to significantly increase its production of oil and natural gas during the last decade—the “Shale Gas Revolution.” As natural gas began to dominate the market with abundant supply and low prices, coal production and consumption have declined. Concurrently, the competitiveness of renewable energy and energy storage has climbed sharply, and analysts expect to see continued reductions in fossil fuel use in the coming decades. Many of these changes have been driven by market forces (i.e., low-cost natural gas and renewables), but current and future policy decisions aimed at tackling climate change concerns and reducing greenhouse gas emissions will also shape the future of the energy sector. This transition to low-carbon fuels has created both opportunities for clean energy technologies and challenges for communities traditionally dependent on fossil fuel-related industries. The power sector’s ongoing shift away from coal has left many coal miners and coal-fired power plant employees unemployed and often unprepared for jobs in other industries, including growing clean energy fields. This primer focuses on the declining coal industry, impacts on communities and workers, opportunities to transition workers who have lost their jobs to clean energy and other related sectors (including hydrogen-oriented jobs), recruitment and training strategies, and available programs and actions to make the shift to a low-carbon economy in a fair, just, and equitable manner by engaging the resources of federal and state governments, as well as the private sector.

01 COAL, LIGNITE, AND PEAT↗

National Solar Jobs Accelerator (Final Technical Report (FTR))

The aptitudes and experiences gained through military service—such as dynamic leadership, teamwork and critical thinking skills, technical specialization, and a mission-completion work ethic—make veterans exceptional candidates for a wide range of solar energy careers. The solar industry offers a highly collaborative and purpose-driven work environment that resonates with service members and veterans looking to rise to their next challenge, and solar employers are eager to tap into this valuable talent pool. From October 2019 - February 2023, The National Solar Jobs Accelerator (publicly the Solar Ready Vets Network TM (SRV Network; SRVN)) enhanced and streamlined options for military service members and veterans to pursue solar training, certification, and employment, while advancing solar employers’ efforts and capacity to invest in military talent as part of a long-term workforce development strategy. The SRVN was led by the Interstate Renewable Energy Council (IREC) in partnership with the Solar Energy Industries Association (SEIA), the US Chamber of Commerce Foundation’s Hiring Our Heroes program (HOH) and the North American Board of Certified Energy Practitioners (NABCEP). Through several direct-impact and indirect, high-impact capacity building initiatives aligned with six key objectives, the SRV Network strengthened solar career pathways, and promoted increased representation of military talent across all levels and sectors of the solar workforce. A work-based learning Corporate Fellowship model connected transitioning service members with on-the-job experience in leadership roles with solar employers nationwide. The project advanced broader veteran recruitment and talent development by expanding GI Bill eligibility and streamlining veterans’ pathways for solar training and credentialing, supported direct connections to jobs with top solar employers, and led coordination among key education and industry partners to advance registered apprenticeships aligned with solar career pathways. To ensure that the project best served the needs of all stakeholders, an Advisory Committee of military-connected solar professionals, solar employers, and training providers met biannually to guide project activities and sustainability plans. The project team engaged the broader “SRV Network” (comprised of over 2,000 veterans, employers and training organizations) through regular newsletters and targeted outreach to share resources, hiring fairs, webinars, and other opportunities for engagement. The work done under this award builds on the previous iterations of the Department of Energy’s Solar Ready Vets ® program. As the solar industry continues to grow rapidly over the next decade, the military community will continue to be a highly valuable source of talent. The relationships established and work accomplished through this project will have an enduring positive impact well beyond the funding period.

14 SOLAR ENERGY↗

Quantifying Uncertainty in HPC Job Queue Time Predictions

High Performance Computing (HPC) has developed at an unprecedented pace in recent decades. This growth has demanded corresponding development in the area of HPC Operational Data Analytics (ODA), which encompasses a wide range of data analysis techniques, ML/AI efforts, tools, and visualizations. Published studies in ODA offer a variety of practical ways to inform HPC users, administrators, procurement managers, and other stakeholders. Uncertainty analysis, however, is rare in the related published literature. For instance, we identify only 1 out of 14 existing studies focused on job queue time prediction that investigates the uncertainty aspect of their proposed predictions. We recognize the utmost importance uncertainty quantification can have in such predictive analytics solutions, with consequences in how users interpret information they receive, and attempt to bridge this gap. With the goal of improving access to such insights, we develop a process for determining upper and lower bounds of the predicted queue times of a regression model at a specified confidence level. Our current research is focused on the uncertainty in predicting job queue times, yet our approach may be employed in predicting other metrics.

HPC↗

Jobs and Economic Development Impact (JEDI) Models

The Jobs and Economic Development Impact (JEDI) models (https://www.nrel.gov/analysis/jedi/) are publicly available, user-friendly tools designed to estimate the economic impacts of both construction and operation of power generation and biofuel plants at a local (usually state) level. Based on project-specific and default inputs (from techno-economic analysis data and NREL's expertise), these models estimate the number of jobs and economic benefits to a region that could reasonably be supported by a particular project. Underlying these calculations is an input-output framework that models the local economy as a network of sectors buying and selling to one-another, creating a multiplier effect. Over the last decade, JEDI has been widely used by both the academic and private sectors, serving as the foundations for multiple peer-reviewed publications and impact analysis reports.

economic impact analysis↗

AIIO: Using Artificial Intelligence for Job-Level and Automatic I/O Performance Bottleneck Diagnosis

Manually diagnosing the I/O performance bottleneck for a single application (hereinafter referred to as the "job level'') is a tedious and error-prone procedure requiring domain scientists to have deep knowledge of complex storage systems. However, existing automatic methods for I/O performance bottleneck diagnosis have one major issue: the granularity of the analysis is at the platform or group level and the diagnosis results cannot be applied to the individual application. To address this issue, we designed and developed a method named "Artificial Intelligence for I/O"(AIIO), which uses AI and its interpretation technology to diagnose I/O performance bottlenecks at the job level automatically. By considering the sparsity of I/O log files, employing multiple AI models for performance prediction, merging diagnosis results across multiple models, and generalizing its performance prediction and diagnosis functions, AIIO can accurately and robustly identify the bottleneck of an even unseen application. Experimental results show that real and unseen applications can use the diagnosis results from AIIO to improve their I/O performance by at most 146 times.

Dong, Bin↗

JobQueue-PG: A Task Queue for Coordinating Varied Tasks Across Multiple HPC Resources and HPC Jobs

The software allows for queueing and dispatch of tasks of small, varied, or uncertain runtimes across multiple HPC jobs, resources, and other computing systems. The software was designed to allow scientists to enqueue, run, and accumulate results from computational experiments in an efficient, manageable manner. For example, the software can be used to enqueue many small computational experiments and run them using several long-running multi-node HPC jobs that may or may not run simultaneously.

Tripp, Charles↗

PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs

It is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead.

Hategan, Mihael↗

Alternative mixed integer linear programming optimization for joint job scheduling and data allocation in grid computing

This paper presents a novel approach to the joint optimization of job scheduling and data allocation in grid computing environments. We formulate this joint optimization problem as a mixed integer quadratically constrained program. To tackle the nonlinearity in the constraint, we alternatively fix a subset of decision variables and optimize the remaining ones via Mixed Integer Linear Programming (MILP). We solve the MILP problem at each iteration via an off-the-shelf MILP solver. Our experimental results show that our method significantly outperforms existing heuristic methods, employing either independent optimization or joint optimization strategies. We have also verified the generalization ability of our method over grid environments with various sizes and its high robustness to the algorithm setting.

97 MATHEMATICS AND COMPUTING↗

International Jobs and Economic Development Impacts Model: Stories of Use and Impact

The International Jobs and Economic Development Impacts (I-JEDI) model is a freely available economic model that estimates gross economic impacts from wind, solar, biopower, and geothermal energy projects around the world. This short publication outlines how the tool has been used by organizations around the world.

economic development↗

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]↗

The Los Angeles 100% Renewable Energy Study (LA100): Chapter 11. Economic Impacts and Jobs

The City of Los Angeles has set ambitious goals to transform its electricity supply, aiming to achieve a 100% renewable energy power system by 2045, along with aggressive electrification targets for buildings and vehicles. To reach these goals, and assess the implications for jobs, electricity rates, the environment, and environmental justice, the Los Angeles City Council passed a series of motions directing the Los Angeles Department of Water and Power (LADWP) to determine the technical feasibility and investment pathways of a 100% renewable energy portfolio standard. The Los Angeles 100% Renewable Energy Study (LA100) is a first-of-its-kind objective, rigorous, and science-based power systems analysis to determine what investments could be made to achieve these goals. The LA100 final report is presented as a collection of 12 chapters and an executive summary, each of which is available as an individual download. This chapter reviews economic impacts, including local net economic impacts and gross workforce impacts.

100% Renewable↗

RCT Continuing Training: RCT Job Coverage and CONOPS

RCT Continuing Training covering all of the necessary steps associated with RCT, and how conduct of operations ties in with this process. The terminal objective of this training is: Given the need to perform radiological job coverage, recognize the requirements to maintain radiological work control in accordance with P121, Radiation Protection, RP-PROG-TP-200, Radiation Protection Manual, and P315, Conduct of Operations.

61 RADIATION PROTECTION AND DOSIMETRY↗