Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “job execution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Automated mesoscale winds derived from GOES multispectral imagery

An automated technique for extracting mesoscale winds from sequences of GOES VISSR image pairs was developed, tested and configured for quasi-real time/research applications on a computing system which gives mesoscale wind estimates at the highest spatial/temporal resolution possible from the VISSR imagery down to a wind vector separation of 10 km. Preprocessing of imagery using IR resampling, VIS edge preserving filtering, and reduced VIS resolution averaging improved height assignments and vector extraction for 10, 15, and 30 min imagery. An objective quality control system provides much greater than 99% accuracy in eliminating questionable wind estimates. Automated winds generally have better spatial coverage and density, and have random error estimates half as large as the manual winds. Dynamical analysis of cloud wind divergence revealed temporally consistent convergence centers on the meso beta scale that are highly correlated with on going and future developing convective storms. The entire system of computer codes was successfully vectorized for execution on an array processor resulting in job turnaround in less than one hour.

Wilson, G. S.↗

Automated mesoscale winds derived from GOES multispectral imagery

An automated technique for extracting mesoscale winds from sequences of GOES visible infrared spin scan radiatiometer (VISSR) image pairs has been developed, tested extensively, and configured for quasi-real time research applications on the Atmospheric Sciences Division's research computing sytem. The entire system of computer codes was successfully vectorized for execution on an array processor resulting in job turnaround in less than 1 hour. An objective quality control system provides much greater than 99 percent accuracy in eliminating questionable wind estimates. Dynamical analysis of cloud wind divergence has revealed temporally consistent convergence centers on the meso-beta scale that are highly correlated with ongoing and future developing convective storms.

Wilson, G. S.↗

Trick Simulation Environment 07

The Trick Simulation Environment is a generic simulation toolkit used for constructing and running simulations. This release includes a Monte Carlo analysis simulation framework and a data analysis package. It produces all auto documentation in XML. Also, the software is capable of inserting a malfunction at any point during the simulation. Trick 07 adds variable server output options and error messaging and is capable of using and manipulating wide characters for international support. Wide character strings are available as a fundamental type for variables processed by Trick. A Trick Monte Carlo simulation uses a statistically generated, or predetermined, set of inputs to iteratively drive the simulation. Also, there is a framework in place for optimization and solution finding where developers may iteratively modify the inputs per run based on some analysis of the outputs. The data analysis package is capable of reading data from external simulation packages such as MATLAB and Octave, as well as the common comma-separated values (CSV) format used by Excel, without the use of external converters. The file formats for MATLAB and Octave were obtained from their documentation sets, and Trick maintains generic file readers for each format. XML tags store the fields in the Trick header comments. For header files, XML tags for structures and enumerations, and the members within are stored in the auto documentation. For source code files, XML tags for each function and the calling arguments are stored in the auto documentation. When a simulation is built, a top level XML file, which includes all of the header and source code XML auto documentation files, is created in the simulation directory. Trick 07 provides an XML to TeX converter. The converter reads in header and source code XML documentation files and converts the data to TeX labels and tables suitable for inclusion in TeX documents. A malfunction insertion capability allows users to override the value of any simulation variable, or call a malfunction job, at any time during the simulation. Users may specify conditions, use the return value of a malfunction trigger job, or manually activate a malfunction. The malfunction action may consist of executing a block of input file statements in an action block, setting simulation variable values, call a malfunction job, or turn on/off simulation jobs.

Lin, Alexander S.↗

Summit Darshan Archival Dataset

Summit Darshan Archival Dataset contains 2021 Summit Darshan log data for 25 applications and is grouped into science domains. The dataset is processed, and all the propriety fields are anonymized. The resultant data is converted into a tabular structure and saved in parquet file format. In this notebook, we demonstrate how to access the data. Data Organization: The data is organized into two directories: Darshan total (`darshan_total`): List all the high levels generated by the `darshan-parser --total` command on `.darshan` files. There is one parquet file for each application. Note: `uid` and `exe` field are masked Darshan detail (`darshan_detail`): This data contains detailed job level log information extracted by command `darshan-parser` on the raw `.darshan` files. The data is sorted by directory hierarchy in the order of `year/month/day (2021/12/07)`. For instance, to get the data for a `job_id` 3819766 of application `App11`, which was executed on `2021-12-07`can be accessed as follows. Note:`uid` and `filename` fields are masked

97 MATHEMATICS AND COMPUTING↗

Accessing protected data by a high-performance computing cluster

A data protection system is provided that allows applications to access protected data in a way that restricts applications from outputting to unauthorized targets any unprotected data derived from the protected data and that ensures that the applications do not have access to a key that allows access to the unprotected data. The data protection system provides a policy server that may execute on a service node of a high performance computing system and a data encryption process that may execute on each compute node that is allocated to an application or batch job. The policy server maintains policies of entities specifying access control for protected data. The data encryption process generates a secure execution environment for an application process and interfaces with the policy server to retrieve keys for decrypting protected data in accordance with a policy, and it decrypts and provides the decrypted data to the application process.

Barnes, Peter↗

Memory protection

Accidental overwriting of files or of memory regions belonging to other programs, browsing of personal files by superusers, Trojan horses, and viruses are examples of breakdowns in workstations and personal computers that would be significantly reduced by memory protection. Memory protection is the capability of an operating system and supporting hardware to delimit segments of memory, to control whether segments can be read from or written into, and to confine accesses of a program to its segments alone. The absence of memory protection in many operating systems today is the result of a bias toward a narrow definition of performance as maximum instruction-execution rate. A broader definition, including the time to get the job done, makes clear that cost of recovery from memory interference errors reduces expected performance. The mechanisms of memory protection are well understood, powerful, efficient, and elegant. They add to performance in the broad sense without reducing instruction execution rate.

Denning, Peter J.↗

Cold Trap Replacement Project Report

This report documents the replacement of the Mechanisms Engineering Test Loop (METL) cold trap. The work involved preparation of the facility to replace the cold trap, removal of the existing welded cold trap from the sodium purification circuit, installation of a new replacement cold trap, completion of associated welds and examinations, restoration of instrumentation and heaters, and controlled return of the cold trap circuit to service. The replacement represented a significant maintenance evolution because the cold trap is an integral welded component of the sodium system. As a result, the work required coordinated control of sodium chemistry, deliberate formation of freeze plugs, inert gas management, precision cutting and welding, and a staged reheating and refill sequence. The activity was executed using procedural controls intended to protect personnel, preserve system cleanliness, and maintain the integrity of the sodium boundary throughout the work. This report provides a narrative summary of the milestone, including the purpose of the work, the pre-job system condition, the major field activities performed, observations made during execution, and the resulting post-work condition of the METL cold trap circuit.

42 ENGINEERING↗

NASTRAN migration to UNIX

COSMIC/NASTRAN, as it is supported and maintained by COSMIC, runs on four main-frame computers - CDC, VAX, IBM and UNIVAC. COSMIC/NASTRAN on other computers, such as CRAY, AMDAHL, PRIME, CONVEX, etc., is available commercially from a number of third party organizations. All these computers, with their own one-of-a-kind operating systems, make NASTRAN machine dependent. The job control language (JCL), the file management, and the program execution procedure of these computers are vastly different, although 95 percent of NASTRAN source code was written in standard ANSI FORTRAN 77. The advantage of the UNIX operating system is that it has no machine boundary. UNIX is becoming widely used in many workstations, mini's, super-PC's, and even some main-frame computers. NASTRAN for the UNIX operating system is definitely the way to go in the future, and makes NASTRAN available to a host of computers, big and small. Since 1985, many NASTRAN improvements and enhancements were made to conform to the ANSI FORTRAN 77 standards. A major UNIX migration effort was incorporated into COSMIC NASTRAN 1990 release. As a pioneer work for the UNIX environment, a version of COSMIC 89 NASTRAN was officially released in October 1989 for DEC ULTRIX VAXstation 3100 (with VMS extensions). A COSMIC 90 NASTRAN version for DEC ULTRIX DECstation 3100 (with RISC) is planned for April 1990 release. Both workstations are UNIX based computers. The COSMIC 90 NASTRAN will be made available on a TK50 tape for the DEC ULTRIX workstations. Previously in 1988, an 88 NASTRAN version was tested successfully on a SiliconGraphics workstation.

Chan, Gordon C.↗

Automated Euler and Navier-Stokes Database Generation for a Glide-Back Booster

The past two decades have seen a sustained increase in the use of high fidelity Computational Fluid Dynamics (CFD) in basic research, aircraft design, and the analysis of post-design issues. As the fidelity of a CFD method increases, the number of cases that can be readily and affordably computed greatly diminishes. However, computer speeds now exceed 2 GHz, hundreds of processors are currently available and more affordable, and advances in parallel CFD algorithms scale more readily with large numbers of processors. All of these factors make it feasible to compute thousands of high fidelity cases. However, there still remains the overwhelming task of monitoring the solution process. This paper presents an approach to automate the CFD solution process. A new software tool, AeroDB, is used to compute thousands of Euler and Navier-Stokes solutions for a 2nd generation glide-back booster in one week. The solution process exploits a common job-submission grid environment, the NASA Information Power Grid (IPG), using 13 computers located at 4 different geographical sites. Process automation and web-based access to a MySql database greatly reduces the user workload, removing much of the tedium and tendency for user input errors. The AeroDB framework is shown. The user submits/deletes jobs, monitors AeroDB's progress, and retrieves data and plots via a web portal. Once a job is in the database, a job launcher uses an IPG resource broker to decide which computers are best suited to run the job. Job/code requirements, the number of CPUs free on a remote system, and queue lengths are some of the parameters the broker takes into account. The Globus software provides secure services for user authentication, remote shell execution, and secure file transfers over an open network. AeroDB automatically decides when a job is completed. Currently, the Cart3D unstructured flow solver is used for the Euler equations, and the Overflow structured overset flow solver is used for the Navier-Stokes equations. Other codes can be readily included into the AeroDB framework.

Chaderjian, Neal M.↗

A microeconomic scheduler for parallel computers

We describe a scheduler based on the microeconomic paradigm for scheduling on-line a set of parallel jobs in a multiprocessor system. In addition to the classical objectives of increasing the system throughput and reducing the response time, we consider fairness in allocating system resources among the users, and providing the user with control over the relative performances of his jobs. We associate with every user a savings account in which he receives money at a constant rate. When a user wants to run a job, he creates an expense account for that job to which he transfers money from his savings account. The job uses the funds in its expense account to obtain the system resources it needs for execution. The share of the system resources allocated to the user is directly related to the rate at which the user receives money; the rate at which the user transfers money into a job expense account controls the job's performance. We prove that starvation is not possible in our model. Simulation results show that our scheduler improves both system and user performances in comparison with two different variable partitioning policies. It is also shown to be effective in guaranteeing fairness and providing control over the performance of jobs.

Stoica, Ion↗

Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure

First-principles fusion plasma simulations are both compute and memory intensive, and CGYRO is no exception. The use of many HPC nodes to fit the problem in the available memory thus results in significant communication overhead, which is hard to avoid for any single simulation. That said, most fusion studies are composed of ensembles of simulations, so we developed a new tool, named XGYRO, that executes a whole ensemble of CGYRO simulations as a single HPC job. By treating the ensemble as a unit, XGYRO can alter the global buffer distribution logic and apply optimizations that are not feasible on any single simulation, but only on the ensemble as a whole. The main saving comes from the sharing of the collisional constant tensor structure, since its values are typically identical between parameter-sweep simulations. This data structure dominates the memory consumption of CGYRO simulations, so distributing it among the whole ensemble results in drastic memory savings for each simulation, which in turn results in overall lower communication overhead.

CGYRO↗

Developing and utilizing an Euler computational method for predicting the airframe/propulsion effects for an aft-mounted turboprop transport. Volume 1: Theory document

An Euler flow solver was developed for predicting the airframe/propulsion integration effects for an aft-mounted turboprop transport. This solver employs a highly efficient multigrid scheme, with a successive mesh-refinement procedure to accelerate the convergence of the solution. A new dissipation model was also implemented to render solutions that are grid insensitive. The propeller power effects are simulated by the actuator disk concept. An embedded flow solution method was developed for predicting the detailed flow characteristics in the local vicinity of an aft-mounted propfan engine in the presence of a flow field induced by a complete aircraft. Results from test case analysis are presented. A user's guide for execution of computer programs, including format of various input files, sample job decks, and sample input files, is provided in an accompanying volume.

Chen, H. C.↗

Competitiveness and Commercialization of Energy Technologies: Supply Chain Deep Dive Assessment

The report “America’s Strategy to Secure the Supply Chain for a Robust Clean Energy Transition” lays out the challenges and opportunities faced by the United States in the energy supply chain as well as the federal government plans to address these challenges and opportunities. It is accompanied by several issue-specific deep dive assessments, including this one, in response to Executive Order 14017 “America’s Supply Chains,” which directs the Secretary of Energy to submit a report on supply chains for the energy sector industrial base. The Executive Order is helping the federal government to build more secure and diverse U.S. supply chains, including energy supply chains. Competitive U.S.-based clean energy manufacturers and rapid commercialization of U.S.-developed technologies are critical to secure energy supply chains, generate high quality jobs, and meet the United States’ national security, energy and climate objectives. The February 2021 “Executive Order on America’s Supply Chains” (E.O. 14017) directs the U.S. Department of Energy (DOE) to evaluate supply chains that encompass the energy industrial base, focusing on technologies that are critical to meet U.S. decarbonization goals by 2050. Understanding and analyzing the end-to-end supply chain through economic analysis is crucial to mitigating risks and identifying opportunities to enhance U.S. competitiveness in the clean energy industry. This insight will allow the Department of Energy (DOE) to leverage its research, development, demonstration and deployment (RDD&D) capabilities to most fully realize the objectives of E.O. 14017.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Using Modules with MPICH-G2 (and "Loose Ends")

A new approach to running complex, distributed MPI jobs using the MPICH-G2 library is described. This approach allows the user to switch between different versions of compilers, system libraries, MPI libraries, etc. via the "module" command. The key idea is a departure from the prescribed "(jobtype=mpi)" approach to running distributed MPI jobs. The new method requires the user to provide a script that will be run as the "executable" with the "(jobtype=single)" RSL attribute. The major advantage of the proposed method is to enable users to decide in their own script what modules, environment, etc. they would like to have in running their job.

Chang, Johnny↗

Simple, Scalable, Script-Based Science Processor (S4P)

The development and deployment of data processing systems to process Earth Observing System (EOS) data has proven to be costly and prone to technical and schedule risk. Integration of science algorithms into a robust operational system has been difficult. The core processing system, based on commercial tools, has demonstrated limitations at the rates needed to produce the several terabytes per day for EOS, primarily due to job management overhead. This has motivated an evolution in the EOS Data Information System toward a more distributed one incorporating Science Investigator-led Processing Systems (SIPS). As part of this evolution, the Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) has developed a simplified processing system to accommodate the increased load expected with the advent of reprocessing and launch of a second satellite. This system, the Simple, Scalable, Script-based Science Processor (S42) may also serve as a resource for future SIPS. The current EOSDIS Core System was designed to be general, resulting in a large, complex mix of commercial and custom software. In contrast, many simpler systems, such as the EROS Data Center AVHRR IKM system, rely on a simple directory structure to drive processing, with directories representing different stages of production. The system passes input data to a directory, and the output data is placed in a "downstream" directory. The GES DAAC's Simple Scalable Script-based Science Processing System is based on the latter concept, but with modifications to allow varied science algorithms and improve portability. It uses a factory assembly-line paradigm: when work orders arrive at a station, an executable is run, and output work orders are sent to downstream stations. The stations are implemented as UNIX directories, while work orders are simple ASCII files. The core S4P infrastructure consists of a Perl program called stationmaster, which detects newly arrived work orders and forks a job to run the appropriate executable (registered in a configuration file for that station). Although S4P is written in Perl, the executables associated with a station can be any program that can be run from the command line, i.e., non-interactively. An S4P instance is typically monitored using a simple Graphical User Interface. However, the reliance of S4P on UNIX files and directories also allows visibility into the state of stations and jobs using standard operating system commands, permitting remote monitor/control over low-bandwidth connections. S4P is being used as the foundation for several small- to medium-size systems for data mining, on-demand subsetting, processing of direct broadcast Moderate Resolution Imaging Spectroradiometer (MODIS) data, and Quick-Response MODIS processing. It has also been used to implement a large-scale system to process MODIS Level 1 and Level 2 Standard Products, which will ultimately process close to 2 TB/day.

Lynnes, Christopher↗

Flow prediction for propfan engine installation effects on transport aircraft at transonic speeds

An Euler-based method for aerodynamic analysis of turboprop transport aircraft at transonic speeds has been developed. In this method, inviscid Euler equations are solved over surface-fitted grids constructed about aircraft configurations. Propeller effects are simulated by specifying sources of momentum and energy on an actuator disc located in place of the propeller. A stripwise boundary layer procedure is included to account for the viscous effects. A preliminary version of an approach to embed the exhaust plume within the global Euler solution has also been developed for more accurate treatment of the exhaust flow. The resulting system of programs is capable of handling wing-body-nacelle-propeller configurations. The propeller disks may be tractors or pushers and may represent single or counterrotation propellers. Results from analyses of three test cases of interest (a wing alone, a wing-body-nacelle model, and a wing-nacelle-endplate model) are presented. A user's manual for executing the system of computer programs with formats of various input files, sample job decks, and sample input files is provided in appendices.

Samant, S. S.↗

DEVELOP’s Approach to Experiential Learning

The NASA DEVELOP Program addresses environmental decision making needs and geoscience workforce development through 10-week feasibility studies that apply Earth observations to environmental issues at hand. The program builds capacity to use geospatial information in both its participants (students, recent graduates, early career professionals, and transitioning career professionals) and partner organizations (federal agencies, state & local governments, non-profits, and private industry). This is accomplished through a structured project execution model that provides opportunities for participants to have autonomy, learn “on the job,” and gain new skillsets for working with remote sensing data. A pipeline of leadership positions enhances opportunities for individuals engaged in the program to get hands-on experience conducting data analyses, communicating their work, leading technical projects, and building their knowledge bank of Earth-observing satellite capabilities. These skillsets and knowledge are then transferred to partners through the projects. This panel contribution will introduce the DEVELOP model, highlight experiences of participants, and share the program’s insights into good practices for effective experiential learning.

Capacity Building↗