Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compute workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Massively Parallel and Portable Genomic Sequence Analysis

Massively Parallel and Portable Genomic Sequence Analysis (mappgene) is a sequencing analysis workflow for high performance computing. It incorporates novel technologies to simplify and accelerate genetics research.

Moon, Josephy↗

CodeScribe Agent

SF-26-086 CodeScribe introduces a structured, multi-stage pipeline that combines deterministic program analysis with LLM-powered translation to enable incremental, testable Fortran-to-C++ migration. First, `code-scribe index` traverses the project directory tree and produces `scribe.yaml` metadata files recording all modules, subroutines, and functions at each level, giving the LLM accurate structural context instead of a hallucinated codebase model. Second, `code-scribe draft` performs the deterministic portion of translation — converting Fortran types to C++ equivalents, replacing `use` statements with `#include` and `using namespace` directives, and detecting constructs requiring special handling — while embedding`scribe-prompt` annotations that guide the LLM through non-trivial cases such as statement-function-to-lambda conversions and `extern "C"` wrapper generation. Third, `code-scribe translate` applies project-specific TOML-based few-shot prompt templates and submits the composed prompt to a pluggable LLM backend (OpenAI, Anthropic, Argonne ARGO, any OpenAI-compatible endpoint, or local Hugging Face checkpoints), producing a C++ source file, a header, and a Fortran-C++ interface file for each translated routine so the codebase compiles and runs correctly throughout the migration. Beyond translation, CodeScribe includes a tool-using coding agent (`code-scribe agent`) with read, bash, edit, and write capabilities, and a bounded loop mode (`code-scribe loop`) that runs repeated stateless agent sessions over a task file with restricted tool access — enabling sustained, auditable software development workflows for broader scientific computing tasks.

Dhruv, Akash [Argonne National Laboratory (ANL), A↗

Two-Step Hyperparameter Optimization Method: Accelerating Hyperparameter Search by Using a Fraction of a Training Dataset

Abstract Hyperparameter optimization (HPO) is an important step in machine learning (ML) model development, but common practices are archaic—primarily relying on manual or grid searches. This is partly because adopting advanced HPO algorithms introduces added complexity to the workflow, leading to longer computation times. This poses a notable challenge to ML applications, as suboptimal hyperparameter selections curtail the potential of ML model performance, ultimately obstructing the full exploitation of ML techniques. In this article, we present a two-step HPO method as a strategic solution to curbing computational demands and wait times, gleaned from practical experiences in applied ML parameterization work. The initial phase involves a preliminary evaluation of hyperparameters on a small subset of the training dataset, followed by a reevaluation of the top-performing candidate models postretraining with the entire training dataset. This two-step HPO method is universally applicable across HPO search algorithms, and we argue it has attractive efficiency gains. As a case study, we present our recent application of the two-step HPO method to the development of neural network emulators for aerosol activation. Although our primary use case is a data-rich limit with many millions of samples, we also find that using up to 0.0025% of the data—a few thousand samples—in the initial step is sufficient to find optimal hyperparameter configurations from much more extensive sampling, achieving up to 135× speedup. The benefits of this method materialize through an assessment of hyperparameters and model performance, revealing the minimal model complexity required to achieve the best performance. The assortment of top-performing models harvested from the HPO process allows us to choose a high-performing model with a low inference cost for efficient use in global climate models (GCMs).

97 MATHEMATICS AND COMPUTING↗

Coupling a Computational Fluid Dynamics Model to a Spacecraft Thermal System Model for the DraMS Instrument Thermal Analysis

The Dragonfly Mass Spectrometer (DraMS) is an instrument on the Dragonfly mission, which will spend 7 years in deep space cruise before landing and operating on the surface of Titan. Vacuum thermal analyses are required for deep space cruise, and convection analyses are required for the Titan surface operations. Model exchanges across multiple thermal teams are needed for all phases of the mission. For DraMS, Thermal Desktop® (TD) has been the main thermal analytical tool of choice due to its capability in modeling complex thermal systems with relatively low computational power and for its availability across thermal teams. However, TD does not have computational fluid dynamics (CFD) capability and struggles to accurately capture complex convective behavior. DraMS has fans operating in tandem and gas flow behaviors are not easily predicted due to its complex flow paths. CFD software, such as Fluent, can model and predict such complex flow behaviors, but CFD models are computationally expensive, and its workflow processes are not tailored towards simulating large and complex systems. Therefore, a coupled modeling approach was chosen for DraMS: A TD model was used for simulating all the conductive, radiative, and source terms, while a Fluent CFD model was added on, as needed, to the TD model to provide the convective boundary conditions using the System Coupling software. The coupling software allows the TD and Fluent models to communicate data and arrive at a co-solved and co-converged solution. Furthermore, Thermal Iso-value Exchange (TIE) method was developed to facilitate and improve the TD-Fluent data exchange process. This paper will discuss the analytical studies that were done to verify the accuracy and usability of the coupled approach and the challenges associated, which lead to the development of the TIE approach. DraMS thermal design and co-solved analysis results will also be discussed.

Heat transfer↗

Coupling a Computational Fluid Dynamics (CFD) Model to a Spacecraft Thermal System Model for the DraMS Instrument Thermal Analysis

The Dragonfly Mass Spectrometer (DraMS) is an instrument on the Dragonfly mission, which will spend 7 years in deep space cruise before landing and operating on the surface of Titan. Vacuum thermal analyses are required for deep space cruise, and convection analyses are required for the Titan surface operations. Model exchanges across multiple thermal teams are needed for all phases of the mission. For DraMS, Thermal Desktop (TD) has been the main thermal analytical tool of choice due to its capability in modeling complex thermal systems with relatively low computational power and for its availability across thermal teams. However, TD does not have computational fluid dynamics (CFD) capability and struggles to accurately capture complex convective behavior. DraMS has fans operating in tandem and gas flow behaviors are not easily predicted due to its complex flow paths. CFD software, such as Fluent, can model and predict such complex flow behaviors, but CFD models are computationally expensive, and its workflow processes are not tailored towards simulating large and complex systems. Therefore, a coupled modeling approach was chosen for DraMS: A TD model was used for simulating all the conductive, radiative, and source terms, while a Fluent CFD model was added on, as needed, to the TD model to provide the convective boundary conditions using the System Coupling software. The coupling software allows the TD and Fluent models to communicate data and arrive at a co-solved and co-converged solution. Furthermore, Thermal Iso-value Exchange (TIE) method was developed to facilitate and improve the TD-Fluent data exchange process. This paper will discuss the analytical studies that were done to verify the accuracy and usability of the coupled approach and the challenges associated, which lead to the development of the TIE approach. DraMS thermal design and co-solved analysis results will also be discussed.

heat transfer↗

Bayesian Optimization for Reactor Design Optimization

This study present a test case in which the Bayesian Optimization method is applied to a simulation-based reactor core design optimization problem. The test case aims to showcase the potential of an automated design optimization algorithm for reactor designs by streamlining the reactor core design workflow, given the high computational cost of simulations. The contributions of this work are threefold. First, the existing HTGR model is converted into a simulation-based design optimization test case by developing a pipeline that enables modification of key design parameters and evaluates design performance based on simulation outputs. Second, Bayesian Optimization is implemented and adapted to demonstrate the feasibility of automatic design optimization for nuclear reactor core. Proposed approach leverages Gaussian Process models to characterize the relationship between design variables and performance metrics, while incorporating novel acquisition functions that balance exploration of the design space with exploitation of promising configurations. This implementation lays the foundation for the future developments of reactor design optimization algorithms.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

rabpro: global watershed boundaries, river elevation profiles, and catchment statistics

River and Basin Profiler (rabpro) is a Python package to delineate watersheds, extract river flowlines and elevation profiles, and compute watershed statistics for any location on the Earth’s surface. As fundamental hydrologically-relevant units of surface area, watersheds are areas of land that drain via aboveground pathways to the same location, or outlet. Delineations of watershed boundaries are typically performed on digital elevation models (DEMs) that represent surface elevations as gridded rasters. Depending on the resolution of the DEM and the size of the watershed, delineation may be very computationally expensive. With this in mind, we designed rabpro to provide user-friendly workflows to manage the complexity and computational expense of watershed calculations given an arbitrary coordinate pair. In addition to basic watershed delineation, rabpro will extract the elevation profile for a watershed’s mainchannel flowline. This enables the computation of river slope, which is a critical parameter in many hydrologic and geomorphologic models. Finally, rabpro provides a user-friendly wrapper around Google Earth Engine’s (GEE) Python API to enable cloud-computing of zonal watershed statistics and/or time-varying forcing data from hundreds of available datasets. Altogether, rabpro provides the ability to automate or semi-automate complex watershed analysis workflows across broad spatial extents.

54 ENVIRONMENTAL SCIENCES↗

A Task Based Approach for Co-Scheduling Ensemble Workloads on Heterogeneous Nodes

Scientific workflows consist of multiple, connected applications, with data and results flowing from one to another in a pipeline. Traditionally, such workflows are executed in sequential order, storing intermediate data in storage disks. Co-scheduling application workflows concurrently on the same compute nodes would greatly reduce the cost of moving data to/from storage and allow real-time analysis of intermediate results. Nevertheless, most parallel programming runtimes do not allow seamless integration of various applications in a scientific workflow, in part due to the complexity of managing data and resources. The situation is even more complicated for heterogeneous systems. In this work we extend the Minos Computing Library (MCL) runtime to accelerate pipe-lined and parallel workloads where multiple applications are running in the same system. MCL’s asynchronous task library and runtime dynamically manages resources to allow co-scheduling of multiple processes sharing heterogeneous resources. In addition, we design a custom ex- tension of the Open Compute Language (OpenCL) to enable multiple processes to share device memory. We enable MCL to coordinate these shared buffers to allow for easy, fast data sharing between applications. Using malleable micro-benchmarks and two application workflows that combine scientific simulation and AI-based analysis, we show that our method outperforms traditional approaches.

Index Terms—Parallel systems, Scheduling and Task ↗

Assessing the Role of Hydrodynamics in Enhancing Height-Above-the-Nearest-Drainage Derived Synthetic Rating Curves: A Comparative Study in the Wu River Basin, Taiwan

The conventional approach to generating synthetic rating curves (SRC) using the Height-Above-the-Nearest-Drainage (HAND) method typically relies on the assumption of uniform flow, such as Manning's equation, to establish stage-discharge ratings. The zero-physics application of the uniform flow equation is insufficient for capturing detailed hydraulic features (e.g., backwater effect) and neglects the hydraulic effects from adjacent channels. This lack of hydrodynamic computation can impact the accuracy and effectiveness of riverine flood risk estimation and management. To reduce this foreseeable error, we introduce the HAND-hd workflow, which integrates sophisticated hydrodynamic computations in the production of HAND-based SRC with hydrodynamic features (SRC hd ). The results indicate that SRC hd demonstrates consistent agreement with both gauge observations and benchmark solutions. Additionally, the comparative analysis suggests that SRC hd provides notable improvements in stage-discharge ratings over conventional HAND-based SRCs, particularly in channels with mild bed gradients, where it reduces water stage prediction errors and percent biases. In steeper channel segments, SRC hd maintains comparable accuracy to conventional methods. The comprehensive evaluation in this study emphasizes the potential discrepancies and inaccuracies associated with the adoption of the uniform flow assumption in the conventional HAND-SRCs and addresses the necessity of including hydrodynamic physics in the application of HAND-based SRC (e.g., inundation map) in channels with mild gradients.

54 ENVIRONMENTAL SCIENCES↗

Optimizing Data Movement for GPU-Based In-Situ Workflow Using GPUDirect RDMA

The extreme-scale computing landscape is increasingly dominated by GPU-accelerated systems. At the same time, in-situ workflows that employ memory-to-memory inter-application data exchanges have emerged as an effective approach for leveraging these extreme-scale systems. In the case of GPUs, GPUDirect RDMA enables third-party devices, such as network interface cards, to access GPU memory directly and has been adopted for intra-application communications across GPUs. In this paper, we present an interoperable framework for GPU-based in-situ workflows that optimizes data movement using GPUDirect RDMA. Specifically, we analyze the characteristics of the possible data movement pathways between GPUs from an in-situ workflow perspective, and design a strategy that maximizes throughput. Furthermore, we implement this approach as an extension of the DataSpaces data staging service, and experimentally evaluate its performance and scalability on a current leadership GPU cluster. The performance results show that the proposed design reduces data-movement time by up to 53% and 40% for the sender and receiver, respectively, and maintains excellent scalability for up to 256 GPUs.

Zhang, Bo↗

Modeling Distributed Computing Infrastructures for HEP Applications

Predicting the performance of various infrastructure design options in complex federated infrastructures with computing sites distributed over a wide area network that support a plethora of users and workflows, such as the Worldwide LHC Computing Grid (WLCG), is not trivial. Due to the complexity and size of these infrastructures, it is not feasible to deploy experimental test-beds at large scales merely for the purpose of comparing and evaluating alternate designs. An alternative is to study the behaviours of these systems using simulation. This approach has been used successfully in the past to identify efficient and practical infrastructure designs for High Energy Physics (HEP). A prominent example is the Monarc simulation framework, which was used to study the initial structure of the WLCG. New simulation capabilities are needed to simulate large-scale heterogeneous computing systems with complex networks, data access and caching patterns. A modern tool to simulate HEP workloads that execute on distributed computing infrastructures based on the SimGrid and WRENCH simulation frameworks is outlined. Studies of its accuracy and scalability are presented using HEP as a case-study. Hypothetical adjustments to prevailing computing architectures in HEP are studied providing insights into the dynamics of a part of the WLCG and candidates for improvements.

Horzela, Maximilian↗

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES↗

Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise

Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. Here, in this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.

Li, Jinyang [George Mason Univ., Fairfax, VA (Unit↗

An end-to-end workflow for executing a classically bootstrapped variational quantum algorithm on an academic quantum computer

Academic quantum computing platforms often face unique challenges in executing quantum workloads due to fragmented software environments and limited engineering support. Unlike commercial ecosystems, academic devices typically evolve without full-stack integration in mind, making it difficult to run complex applications—such as variational quantum algorithms (VQA)—reliably and efficiently. Issues such as incompatible software layers and lack of automated job management significantly increase the overhead of theory-experiment collaboration. To address these challenges, we develop a modular, end-to-end workflow that decouples application-layer code from low-level hardware control, automates circuit submission and result collection, and supports fine-grained circuit-level job scheduling and recovery. The architecture employs a dual-end application programming interface (API) design, enabling robust operation across unstable or resource-constrained hardware backends. For practical use, the framework is lightweight and user-friendly, allowing rapid prototyping of full-stack workflows using basic Python tools. We validate this workflow on a high-fidelity trapped-ion quantum computer by demonstrating a variational quantum eigensolver (VQE) experiment with a classically bootstrapped ansatz initialization technique. The system successfully executed over 60,000 circuits across multiple molecular test cases with minimal human intervention, highlighting the framework’s effectiveness in enabling reproducible, resilient quantum experimentation in academic settings.

Clifford↗

Cloud Services Enable Efficient AI-Guided Simulation Workflows across Heterogeneous Resources

Applications which fuse machine learning and simulation are rarely best served by a single computing resource. Highly parallel simulation codes are best deployed on super- computers, while AI tasks used to decide which simulations to perform may be best suited to specialized accelerators. Here we present a Function-as-a-Service (FaaS) system for executing complex, distributed computational campaigns that achieves performance parity with conventional workflow systems without the complexities of secure network connections between compute providers. One innovation enabling high performance is a subsystem that directly moves task data between sites, separate from the cloud-hosted FaaS system used to distribute task instructions. We also introduce a flexible scheduling system that allows us access factor of 2 trade offs between the amount of resources required to solve a problem at each compute site. We anticipate that this system will upgrade multi-site applications from demonstration projects to routine practice in computational science.

Ward, Logan↗

MISPR : an open-source package for high-throughput multiscale molecular simulations

Computational tools provide a unique opportunity to study and design optimal materials by enhancing our ability to comprehend the connections between their atomistic structure and functional properties. However, designing materials with tailored functionalities is complicated due to the necessity to integrate various computational-chemistry software (not necessarily compatible with one another), the heterogeneous nature of the generated data, and the need to explore vast chemical and parameter spaces. The latter is especially important to avoid bias in scattered data points-based models and derive statistical trends only accessible by systematic datasets. Here, we introduce a robust high-throughput multi-scale computational infrastructure coined MISPR (Materials Informatics for Structure–Property Relationships) that seamlessly integrates classical molecular dynamics (MD) simulations with density functional theory (DFT). By enabling high-performance data analytics and coupling between different methods and scales, MISPR addresses critical challenges arising from the needs of automated workflow management and data provenance recording. The major features of MISPR include automated DFT and MD simulations, error handling, derivation of molecular and ensemble properties, and creation of output databases that organize results from individual calculations to enable reproducibility and transparency. In this work, we describe fully automated DFT workflows implemented in MISPR to compute various properties such as nuclear magnetic resonance chemical shift, binding energy, bond dissociation energy, and redox potential with support for multiple methods such as electron transfer and proton-coupled electron transfer reactions. The infrastructure also enables the characterization of large-scale ensemble properties by providing MD workflows that calculate a wide range of structural and dynamical properties in liquid solutions. MISPR employs the methodologies of materials informatics to facilitate understanding and prediction of phenomenological structure–property relationships, which are crucial to designing novel optimal materials for numerous scientific applications and engineering technologies.

36 MATERIALS SCIENCE↗

PROcess Based Diagnostics PROBE

Many of the aspects of the climate system that are of the greatest interest (e.g., the sensitivity of the system to external forcings) are emergent properties that arise via the complex interplay between disparate processes. This is also true for climate models most diagnostics are not a function of an isolated portion of source code, but rather are affected by multiple components and procedures. Thus any model-observation mismatch is hard to attribute to any specific piece of code or imperfection in a specific model assumption. An alternative approach is to identify diagnostics that are more closely tied to specific processes -- implying that if a mismatch is found, it should be much easier to identify and address specific algorithmic choices that will improve the simulation. However, this approach requires looking at model output and observational data in a more sophisticated way than the more traditional production of monthly or annual mean quantities. The data must instead be filtered in time and space for examples of the specific process being targeted.We are developing a data analysis environment called PROcess-Based Explorer (PROBE) that seeks to enable efficient and systematic computation of process-based diagnostics on very large sets of data. In this environment, investigators can define arbitrarily complex filters and then seamlessly perform computations in parallel on the filtered output from their model. The same analysis can be performed on additional related data sets (e.g., reanalyses) thereby enabling routine comparisons between model and observational data. PROBE also incorporates workflow technology to automatically update computed diagnostics for subsequent executions of a model. In this presentation, we will discuss the design and current status of PROBE as well as share results from some preliminary use cases.

PROBE↗

Performance Analysis of Cloud Computing Architectures Using Discrete Event Simulation

Cloud computing offers the economic benefit of on-demand resource allocation to meet changing enterprise computing needs. However, the flexibility of cloud computing is disadvantaged when compared to traditional hosting in providing predictable application and service performance. Cloud computing relies on resource scheduling in a virtualized network-centric server environment, which makes static performance analysis infeasible. We developed a discrete event simulation model to evaluate the overall effectiveness of organizations in executing their workflow in traditional and cloud computing architectures. The two part model framework characterizes both the demand using a probability distribution for each type of service request as well as enterprise computing resource constraints. Our simulations provide quantitative analysis to design and provision computing architectures that maximize overall mission effectiveness. We share our analysis of key resource constraints in cloud computing architectures and findings on the appropriateness of cloud computing in various applications.

Stocker, John C.↗