Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “job execution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A user-oriented synthetic workload generator

A user oriented synthetic workload generator that simulates users' file access behavior based on real workload characterization is described. The model for this workload generator is user oriented and job specific, represents file I/O operations at the system call level, allows general distributions for the usage measures, and assumes independence in the file I/O operation stream. The workload generator consists of three parts which handle specification of distributions, creation of an initial file system, and selection and execution of file I/O operations. Experiments on SUN NFS are shown to demonstrate the usage of the workload generator.

Kao, Wei-Lun↗

Execute BEE workflows on private cloud infrastructure-2.3.6.01 - LANL ATDM ST / STNS01-22 Milestone Completion Documentation (BEE-FY21 P6-2) [Slides]

This work involves the creation of the Cloud Launcher, a new subcomponent of BEE, and the extension of the BEETaskManager to run on Cloud systems. BEE will be able to interact with the Google Compute Engine and OpenStack cloud APIs to set up simple Cloud clusters for launching HPC job scripts. BEE will use existing functionality to launch jobs that previously could only be launched on HPC systems. The BEETaskManager will handle launching tasks on the Cloud cluster.

97 MATHEMATICS AND COMPUTING↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗

Combining Structural and Substructural Mathematical Models

Automation reduces input length and potential data-entry errors. Matrix Automated Reduction and Coupling (MARC) program used for combining NASTRAN substructural models with primary structural model MARC also constructs job-control language (JCL) stream for NASTRAN batch job that utilizes previously-written user library of dynamic models. Minimizes lengthy input and reduces potential data-entry errors. MARC procedure used in assembling Space Shuttle orbiter dynamic models since 1983 and reduced NASTRAN modeling input time by as much as 50 percent. MARC program written in FORTRAN IV for interactive execution.

Choa, V. K.↗

The ALEXIS data processing package: An IDL based system

The Array of Low Energy X-ray Imaging Sensors (ALEXIS) experiment consists of a mini-satellite containing six wide angle EUV/ultrasoft x-ray telescopes. Its purpose is to map out the sky in three narrow (approximately 5 percent) bandpasses around 66, 71, and 93 eV. The 66 and 71 eV bandpasses are centered on intense Fe emission lines which are characteristic of million degree plasmas such as the one thought to produce the soft x-ray background. The 93 eV bandpass is not near any strong emission lines and is more sensitive to continuum sources. The mission will be launched on the Pegasus Air Launched Vehicle in the second half of 1992 into a 400-nautical-mile, high inclination orbit and will be controlled entirely from a small ground station located at Los Alamos. The project is a collaborative effort between Los Alamos National Laboratory, Sandia National Laboratory, and the University of California-Berkeley Space Sciences Laboratory. The six telescopes are arranged in three pairs. As the satellite spins twice a minute they scan the entire anti-solar hemisphere. Each f/1 telescope consists of a spherical, multilayer-coated mirror with a curved, microchannel plate detector located at the prime focus. The multilayer coatings determine the bandpasses of the telescopes. The field of view of each telescope is 30 degrees with a spatial resolution of 0.5 degree, limited by spherical aberration. The data processing requirements for ALEXIS are large. Each event is one of the six telescopes is telemetered to the ground with its time of arrival and position on the detector. This information must be folded with the aspect solution for the satellite to reconstruct the direction on the sky from which the photon came. Because of the way the six telescopes scan the sky, the effective exposure calculation is also very computationally intensive. ALEXIS may generate up to 100 megabytes of raw data per day, which are converted into a gigabyte per day of processed data. While the processing job for ALEXIS is sizable, the programming staff is small. To maximize programming efficiency, and to make the best use of tools available in the public domain, we chose IDL as our software development platform. IDL was used from the start of instrument development through flight. We use IDL as a top-level executive for the processing tasks (replacing Unix shell scripts), as a device independent graphics engine, as a database manager, and as a final data manipulator. IDL routines spawn special purpose C programs to perform detailed telemetry deconvolution and other specialized functions. We discuss the use of IDL and C within the processing and archiving strategy for the ALEXIS data anlaysis system as implemented on a SPARCstation platform. We also show results from our End-to-End software simulation capability as processed by our analysis codes.

Bloch, J. J.↗

Software fault tolerance in computer operating systems

This chapter provides data and analysis of the dependability and fault tolerance for three operating systems: the Tandem/GUARDIAN fault-tolerant system, the VAX/VMS distributed system, and the IBM/MVS system. Based on measurements from these systems, basic software error characteristics are investigated. Fault tolerance in operating systems resulting from the use of process pairs and recovery routines is evaluated. Two levels of models are developed to analyze error and recovery processes inside an operating system and interactions among multiple instances of an operating system running in a distributed environment. The measurements show that the use of process pairs in Tandem systems, which was originally intended for tolerating hardware faults, allows the system to tolerate about 70% of defects in system software that result in processor failures. The loose coupling between processors which results in the backup execution (the processor state and the sequence of events occurring) being different from the original execution is a major reason for the measured software fault tolerance. The IBM/MVS system fault tolerance almost doubles when recovery routines are provided, in comparison to the case in which no recovery routines are available. However, even when recovery routines are provided, there is almost a 50% chance of system failure when critical system jobs are involved.

Iyer, Ravishankar K.↗

A cost-effective intelligent robotic system with dual-arm dexterous coordination and real-time vision

Dexterous coordination of manipulators based on the use of redundant degrees of freedom, multiple sensors, and built-in robot intelligence represents a critical breakthrough in development of advanced manufacturing technology. A cost-effective approach for achieving this new generation of robotics has been made possible by the unprecedented growth of the latest microcomputer and network systems. The resulting flexible automation offers the opportunity to improve the product quality, increase the reliability of the manufacturing process, and augment the production procedures for optimizing the utilization of the robotic system. Moreover, the Advanced Robotic System (ARS) is modular in design and can be upgraded by closely following technological advancements as they occur in various fields. This approach to manufacturing automation enhances the financial justification and ensures the long-term profitability and most efficient implementation of robotic technology. The new system also addresses a broad spectrum of manufacturing demand and has the potential to address both complex jobs as well as highly labor-intensive tasks. The ARS prototype employs the decomposed optimization technique in spatial planning. This technique is implemented to the framework of the sensor-actuator network to establish the general-purpose geometric reasoning system. The development computer system is a multiple microcomputer network system, which provides the architecture for executing the modular network computing algorithms. The knowledge-based approach used in both the robot vision subsystem and the manipulation control subsystems results in the real-time image processing vision-based capability. The vision-based task environment analysis capability and the responsive motion capability are under the command of the local intelligence centers. An array of ultrasonic, proximity, and optoelectronic sensors is used for path planning. The ARS currently has 18 degrees of freedom made up by two articulated arms, one movable robot head, and two charged coupled device (CCD) cameras for producing the stereoscopic views, and articulated cylindrical-type lower body, and an optional mobile base. A functional prototype is demonstrated.

Marzwell, Neville I.↗

Carbon Capture from ArcelorMittal Hot Briquetted Iron Plant Using Air Liquide Cryocap™ FG Technology – FEED Study

The process of steel production is energy and carbon intensive with global average energy consumption of 5.5 MWh/tonne of steel and CO2 emission intensity of 1.83 tonne CO2/tonne of steel. The steel making process has inherent CO2 emissions from mineral conversion and is considered major contributors to the global carbon emissions. The steel industry is responsible for 8% of global carbon emissions. The main objective of this research project is to execute and complete a front-end engineering and design (FEED) study for a commercial-scale, carbon capture project that separates 95% of the total CO2 emissions at the ArcelorMittal’s Hot Briquetted Iron (HBI) plant in Portland, TX (Figure 1). The HBI is an ore-based metallic that is used as high-grade feedstock for high-quality steel via an Electric Arc Furnace (EAF) route. The HBI plant produces 2.0 million metric tonnes of high-quality HBI and emits approximately 1 million tonnes CO2/yr. The capture system is a Pressure Swing Adsorption (PSA) system assisted Cryocap™ FG technology (Figure 2). The captured CO2 will be pipeline grade and will be geologically stored in a facility within 10 miles of the CO2 source. The Host Site location in Corpus Christi, TX, is near hydrocarbon processing facilities and near Environmental Justice (EJ) and Qualified Opportunity Zone (QOZ) communities. Due to the location of the Host Site, the retrofit project offers the ability to demonstrate how a workforce focused on the fossil energy sector can be redirected to the clean- energy sector. The Air Liquide Cryocap™ capture technology is a proven technology and has been extensively examined for large industrial applications. It has been shown to be applicable to a variety of industrial applications including the steel industry. Cryocap™ FG (specific setup for Flue Gas application) consists of a Pressure Swing Adsorption (PSA) unit coupled with a Cryogenic System. The PSA pre-concentrates the CO2 from the flue gas, while the cryogenic unit enables the CO2 purity to be increased to the desired level. The scope of this study incorporates completing FEED study of the CO2 capture system which includes point-source CO2 capture and balance-of-plant; Business Case Analysis (BCA) outlining the current and projected volumes of the steel plant’s point sources of CO2 and the potential utilization of tax credits, including its projected revenue and duration; Life Cycle Analysis (LCA); Environmental Justice Analysis; Economic Revitalization and Job Creation Outcomes Analysis; and Workforce Readiness Plan. The plant design work was divided into two components: Inside Battery Limits (ISBL) and Outside Battery Limits (OSBL). The ISBL focuses on the capture system, while the OSBL focuses on the utility feeds and ducting from the plant to the capture system. Various design and engineering deliverables will be developed to define commodity quantities, equipment specifications, and labour effort required to execute the project. These FEED study deliverables will be prepared with the intent to develop an overall project capital cost estimate consistent with an AACE Class 3 estimate. The modular approach for the Cryocap™ FG that is being designed for this study integrates compression, PSA, and cryogenic “bricks” to achieve the desired CO2 capture rates. This carbon capture system integrates easily with the existing plant, thus reducing project costs and risks. It is also capable of managing impurities such as nitrogen oxides (NOx), sulfur oxides (SOx), mercury, hydrocarbons, and particulate matter. The capture system has a smaller footprint than amine-based systems. The two-step process uses PSA to preconcentrate the CO2 in the feedstream and then uses the cryogenic portion to purify and compress the resulting high purity CO2 product. This combination of purification and compression (i.e., process intensification) significantly reduces the CAPEX associated with use of a separate compressor commonly utilized for amine solvent-based systems. Successful completion of the FEED study will provide DOE with a detailed understanding of the costs for scaling up this proven capture technology for commercial applications at industrial facilities.

42 ENGINEERING↗

Heterogeneous Reconstruction of Tracks and Primary Vertices With the CMS Pixel Tracker

The High-Luminosity upgrade of the Large Hadron Collider (LHC) will see the accelerator reach an instantaneous luminosity of 7 × 10 34 cm −2 s −1 with an average pileup of 200 proton-proton collisions. These conditions will pose an unprecedented challenge to the online and offline reconstruction software developed by the experiments. The computational complexity will exceed by far the expected increase in processing power for conventional CPUs, demanding an alternative approach. Industry and High-Performance Computing (HPC) centers are successfully using heterogeneous computing platforms to achieve higher throughput and better energy efficiency by matching each job to the most appropriate architecture. In this paper we will describe the results of a heterogeneous implementation of pixel tracks and vertices reconstruction chain on Graphics Processing Units (GPUs). The framework has been designed and developed to be integrated in the CMS reconstruction software, CMSSW. The speed up achieved by leveraging GPUs allows for more complex algorithms to be executed, obtaining better physics output and a higher throughput.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Clippy

Clippy (CLI + PYthon) is a Python language interface to HPC resources. Precompiled binaries that execute on HPC systems are exposed as methods to a dynamically-created Clippy Python object, where they present a familiar interface to researchers, data scientists, and others. Clippy allows these users to interact with HPC resources in an easy, straightforward environment - at the REPL, for example, or within a notebook - without the need to learn complex HPC behavior and arcane job submission commands.

Bromberger, SethA.↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗

BEE - FY20 P6-3: Release BEEWorkflowManager, BEETaskManager, and client application 2.3.6.01 – LANL ATDM ST / STNS01-4 P6 Milestone Completion Documentation

Release BEEWorkflowManager, BEETaskManager, and client software. The BEEWorkflowManager daemon runs on the HPC cluster login node. It accepts workflows submitted by the BEE client. These workflows are specified using the Common Workflow Language (CWL) standard. The BEEWorkflowManager loads workflows into the Neo4j graph database to create the workflow directed acyclic graph (DAG), and submits the workflow tasks to the BEETaskManager for execution. The BEEWorkflowManager records the state of the workflow and its tasks, and communicates this state to the BEE client. The BEEWorkflowManager will start, pause, and cancel a running workflow and its tasks at the command of the BEE client. The BEETaskManager daemon runs on the HPC cluster login node. It accepts tasks from the BEEWorkflowManager, turns those tasks into HPC resource manager jobs (e.g. a slurm job script), and submits the job to the cluster resource manager. The BEETaskManager then tracks the status of the job (pending, running, complete) and updates the BEEWorkflowManager. The BEETaskManager will also cancel a queued or running job when commanded to do so by the BEEWorkflowManager. The first release of the BEETaskManager will support the Slurm resource manager and the Charliecloud linux container runtime.

97 MATHEMATICS AND COMPUTING↗

Batching System for Superior Service

Veridian's Portable Batch System (PBS) was the recipient of the 1997 NASA Space Act Award for outstanding software. A batch system is a set of processes for managing queues and jobs. Without a batch system, it is difficult to manage the workload of a computer system. By bundling the enterprise's computing resources, the PBS technology offers users a single coherent interface, resulting in efficient management of the batch services. Users choose which information to package into "containers" for system-wide use. PBS also provides detailed system usage data, a procedure not easily executed without this software. PBS operates on networked, multi-platform UNIX environments. Veridian's new version, PBS Pro,TM has additional features and enhancements, including support for additional operating systems. Veridian distributes the original version of PBS as Open Source software via the PBS website. Customers can register and download the software at no cost. PBS Pro is also available via the web and offers additional features such as increased stability, reliability, and fault tolerance.A company using PBS can expect a significant increase in the effective management of its computing resources. Tangible benefits include increased utilization of costly resources and enhanced understanding of computational requirements and user needs.

Source record↗

Profiles of upcoming HPC Applications and their Impact on Reservation Strategies

With the expected convergence between HPC, BigData and AI, new applications with different profiles are coming to HPC infrastructures. Here, we aim at better understanding the features and needs of these applications in order to be able to run them efficiently on HPC platforms. The approach followed is bottom-up: we study thoroughly an emerging application from the neuroscience community (SLANT) to understand its behavior. Based on these observations, we derive a generic, yet simple, application model (namely, a linear sequence of stochastic jobs). We expect this model to be representative for a large set of upcoming applications that require the computational power of HPC clusters without fitting the typical behavior of large-scale traditional applications. In a second step, we show how one can manipulate this generic model in a scheduling framework. Specifically we consider the problem of making reservations (both time and memory) for an execution on an HPC platform. We derive solutions using the model of the first step of this work. We experimentally show the robustness of the model, even with very few data or with another application, to generate the model, and provide performance gains with regards to standard and more recent approaches used in the neuroscience community.

97 MATHEMATICS AND COMPUTING↗

Using a Cray Y-MP as an array processor for a RISC Workstation

As microprocessors increase in power, the economics of centralized computing has changed dramatically. At the beginning of the 1980's, mainframes and super computers were often considered to be cost-effective machines for scalar computing. Today, microprocessor-based RISC (reduced-instruction-set computer) systems have displaced many uses of mainframes and supercomputers. Supercomputers are still cost competitive when processing jobs that require both large memory size and high memory bandwidth. One such application is array processing. Certain numerical operations are appropriate to use in a Remote Procedure Call (RPC)-based environment. Matrix multiplication is an example of an operation that can have a sufficient number of arithmetic operations to amortize the cost of an RPC call. An experiment which demonstrates that matrix multiplication can be executed remotely on a large system to speed the execution over that experienced on a workstation is described.

Lamaster, Hugh↗

Repurposing of the Run 2 CMS High Level Trigger Infrastructure as a Cloud Resource for Offline Computing

The former CMS Run 2 High Level Trigger (HLT) farm is one of the largest contributors to CMS compute resources, providing about 25k job slots for offline computing. This CPU farm was initially employed as an opportunistic resource, exploited during inter-fill periods, in the LHC Run 2. Since then, it has become a nearly transparent extension of the CMS capacity at CERN, being located on-site at the LHC interaction point 5 (P5), where the CMS detector is installed. This resource has been configured to support the execution of critical CMS tasks, such as prompt detector data reconstruction. It can therefore be used in combination with the dedicated Tier 0 capacity at CERN, in order to process and absorb peaks in the stream of data coming from the CMS detector. The initial configuration for this resource, based on statically configured VMs, provided the required level of functionality. However, regular operations of this cluster revealed certain limitations compared to the resource provisioning and use model employed in the case of WLCG sites. A new configuration, based on a vacuum-like model, has been implemented for this resource in order to solve the detected shortcomings. This paper reports about this redeployment work on the permanent cloud for an enhanced support to CMS offline computing, comparing the former and new models’ respective functionalities, along with the commissioning effort for the new setup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability

The Cray XK7 Titan was the top supercomputer system in the world for a long time and remained critically important throughout its nearly seven year life. It was an interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 years of GPU lifetimes during Titan’s 6-year-long productive period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the cooling architecture and job scheduling. We describe the history, data collection, cleaning, and analysis and give recommendations for future supercomputing systems. We make the data and our analysis codes publicly available.

Ostrouchov, George↗