Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modeling workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

Automated ICRF heating surrogate modeling via machine learning

This work introduces automated machine learning workflows that address critical bottlenecks in surrogate model development for Ion Cyclotron Range of Frequencies (ICRF) heating applications. The automated framework includes data analysis tools that transform raw datasets into actionable insights in seconds, replacing weeks of manual exploratory effort and ensuring consistent, reproducible dataset characterization. By integrating advanced hyperparameter optimization (HPO) methods including Bayesian optimization via BoTorch and Tree-structured Parzen Estimators (TPE), the framework significantly reduces model development time from weeks to hours, decreasing computational cost and required expertise, while enabling high-accuracy surrogate models. Compared to traditional hyperparameter scanning (HPS) techniques such as methodical, randomized, and grid searches, HPO methods achieve superior convergence and predictive performance, even when compared to already well-tuned reference models. On NSTX High Harmonic Fast Wave (HHFW) heating datasets, both Random Forest Regressor (RFR) and neural network surrogates demonstrate improved accuracy, achieving R 2 values beyond 0.97 and 0.98, respectively. The results show that while HPO gains are modest for robust architectures like RFR, they become essential for more sensitive models such as neural networks, highlighting the trade-offs across optimization strategies. Through automated workflows that eliminate manual hyperparameter tuning and require minimal ML expertise, this work enables widespread adoption of high-fidelity surrogate models across the fusion community for real-time plasma control, uncertainty quantification, rapid experimental scenario development, and integrated system optimization.

Sanchez-Villar, Alvaro [Princeton Plasma Physics L↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Exploration of multifidelity UQ sampling strategies for computer network applications.

Network modeling is a powerful tool to enable rapid analysis of complex systems that can be challenging to study directly using physical testing. Here, two approaches are considered: emulation and simulation. The former runs real software on virtualized hardware, while the latter mimics the behavior of network components and their interactions in software. Although emulation provides an accurate representation of physical networks, this approach alone cannot guarantee the characterization of the system under realistic operative conditions. Operative conditions for physical networks are often characterized by intrinsic variability (payload size, packet latency, etc.) or a lack of precise knowledge regarding the network configuration (bandwidth, delays, etc.); therefore uncertainty quantification (UQ) strategies should be also employed. UQ strategies require multiple evaluations of the system with a number of evaluation instances that roughly increases with the problem dimensionality, i.e., the number of uncertain parameters. It follows that a typical UQ workflow for network modeling based on emulation can easily become unattainable due to its prohibitive computational cost. In this paper, a multifidelity sampling approach is discussed and applied to network modeling problems. The main idea is to optimally fuse information coming from simulations, which are a low-fidelity version of the emulation problem of interest, in order to decrease the estimator variance. By reducing the estimator variance in a sampling approach it is usually possible to obtain more reliable statistics and therefore a more reliable system characterization. Several network problems of increasing difficulty are presented. For each of them, the performance of the multifidelity estimator is compared with respect to the single fidelity counterpart, namely, Monte Carlo sampling. For all the test problems studied in this work, the multifidelity estimator demonstrated an increased efficiency with respect to MC.

97 MATHEMATICS AND COMPUTING↗

Providing performance portable numerics for Intel GPUs

Summary With discrete Intel GPUs entering the high‐performance computing landscape, there is an urgent need for production‐ready software stacks for these platforms. In this article, we report how we enable the Ginkgo math library to execute on Intel GPUs by developing a kernel backed based on the DPC++ programming environment. We discuss conceptual differences between the CUDA and DPC++ programming models and describe workflows for simplified code conversion. We evaluate the performance of basic and advanced sparse linear algebra routines available in Ginkgo's DPC++ backend in the hardware‐specific performance bounds and compare against routines providing the same functionality that ship with Intel's oneMKL vendor library.

97 MATHEMATICS AND COMPUTING↗

Artificial intelligence–powered biofoundries for protein engineering and metabolic engineering

Synthetic biology is rapidly evolving through the integration of artificial intelligence (AI) and automated biofoundries. This convergence accelerates the design–build–test–learn cycle, shifting protein engineering and metabolic engineering from labor-intensive manual experimentation to autonomous experimentation. This review summarizes recent advances in workflow development, AI models, and their integration with biofoundries for automated or autonomous protein engineering and metabolic engineering. Particularly, we highlight the potential of AI-powered biofoundries for accelerated scientific discovery and innovation in synthetic biology.

Chen, Junyu [Univ. of Illinois at Urbana-Champaign↗

Deep generative learning of magnetic frustration in artificial spin ice from magnetic force microscopy images

Increasingly large datasets of microscopic images with nanoscale resolution facilitate the development of machine learning methods to identify and analyze subtle physical phenomena embedded within the images. In this work, microscopic images of honeycomb lattice spin-ice samples serve as datasets from which we automate the calculation of net magnetic moments and directional orientations of spin-ice configurations. In the first stage of our workflow, machine learning models are trained to accurately predict magnetic moments and directions within spin-ice structures. Variational Autoencoders (VAEs), an emergent unsupervised deep learning technique, are employed to generate high-quality synthetic magnetic force microscopy (MFM) images and extract latent feature representations, thereby reducing experimental and segmentation errors. The second stage of proposed methodology enables precise identification and prediction of frustrated vertices and nanomagnetic segments, effectively correlating structural and functional aspects of microscopic images. This facilitates the design of optimized spin-ice configurations with controlled frustration patterns, enabling potential on-demand synthesis.

36 MATERIALS SCIENCE↗

Market optimization and technoeconomic analysis of hydrogen-electricity coproduction systems

Decarbonization efforts across North America, Europe, and beyond rely on variable renewable energy sources such as wind and solar, as well as alternative fuels, such as hydrogen, to support the sustainable energy transition. These advancements have prompted a need for more flexibility in the electric grid to complement non-dispatchable energy sources and increased demand from electrification. Integrated energy systems are well suited to provide this flexibility, but conventional technoeconomic modeling paradigms neglect the time-varying dynamic nature of the grid and thus undervalue resource flexibility. In this work, we develop a computational optimization framework for dynamic market-based technoeconomic comparison of integrated energy systems that coproduce low-carbon electricity and hydrogen (e.g., solid oxide fuel cells, solid oxide electrolysis) against technologies that only produce electricity (e.g., natural gas combined cycle with carbon capture) or only produce hydrogen. Our framework starts with rigorous physics-based process models, built in the open-source Institute for the Design of Advanced Energy Systems (IDAES) modeling and optimization platform, for six energy process concepts. Using these rigorous models and a workflow to optimally design each technology, the framework is shown to be capable of evaluating new and emerging technologies in varying energy markets under a plethora of future scenarios (i.e., renewables penetration, carbon tax, etc.). Ultimately, our framework finds that solid oxide fuel cell-based coproduction systems achieve positive profits for 85% of the analyzed market scenarios. From these market optimization results, we use multivariate linear regression (R 2 values up to 0.99) to determine which electricity price statistics are most significant to predict the optimized annual profit of each system. The proposed framework provides a powerful tool for directly comparing flexible, multi-product energy process concepts to help discern optimal technology and integration options.

08 HYDROGEN↗

The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism

Here, we present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy.

97 MATHEMATICS AND COMPUTING↗

The Value of Hyperparameter Optimization in Phase-Picking Neural Networks

The effectiveness of neural networks for picking seismic phase arrival times has been demonstrated through several case studies, and seismic monitoring programs are starting to adopt the technology into their workflows. However, published models were designed and trained using rather arbitrary choices of hyperparameters, limiting their performance. In this study, we use phase picks from both routine and template-matching analyses from multiple regions (Ridgecrest, California; Kilauea, Hawaii; Yellowstone, Wyoming–Montana–Idaho) to test a hyperparameter optimization scheme for phase-picking neural networks and to evaluate their performance. We show that a published model, namely PhaseNet (Zhu and Beroza, 2019), can be simplified and improved with reasonable effort and there are preferred choices of hyperparameters that increase the performance. We also show that models optimized based on the arrival times reported in routine event catalogs consistently perform well when picking arrival times of smaller events, which is crucial for many tasks from microseismicity to explosion monitoring.

58 GEOSCIENCES↗

Basin-Scale Structural Features Database

The Basin-Scale Structural Features database provides spatial datasets of faults, fractures, folds, and earthquakes compiled from public, authoritative sources (e.g., U.S. Geological Survey and State Geological Surveys) and aggregated into derivative forms to support subsurface assessments. Recognizing that characterizing basin-scale structural features requires interpreting data that are often ambiguous or lack key information, the source data were evaluated using a knowledge-data framework and geospatial fuzzy logic method (Justman et al., 2020) to represent both measured (observed) and predicted (inferred or potential) structural features as derivative datasets. This workflow employs conceptual models for known structural features and predicted structural features, incorporating geospatial data to estimate potential, even with limited data. The aim is to aid and support an understanding of basin-scale features and identify potential gaps in data and knowledge. As of 4/30/2025, the database includes resources for nine sedimentary basins: Appalachian, Denver, U.S. Gulf Coast, Illinois, Michigan, Permian, Sacramento, San Joquin and Williston. The database is organized by basin and then data category: 1) Faults, fractures, folds, 2) Earthquakes, 3) Topographic, 4) Structural contours and isopachs, 5) Geophysical, and 6) Structural feature density assessment maps.

basin scale↗

System Modeling Frameworks for Wind Turbines and Plants: Review and Requirements Specifications

System modeling frameworks for wind turbines and plants are used by research groups and industry to design wind energy systems that take into account key trade-offs across performance, cost, and reliability at both the turbine and plant level. The frameworks are exercised using a variety of multi-disciplinary design, analysis and optimization (MDAO) methods. To improve inter-operability and foster collaboration, this report proposes a classification system for the frameworks along dimensions of model fidelity and scope. The classification system is first motivated with reviews the state-of-the-art in the development of software frameworks for integrated wind turbine and plant simulation. Within each major wind turbine and power plant subsystem, a matrix is developed for the disciplines used and the fidelity levels with which each discipline can be modeled. The existing frameworks are then classified according to the matrix. Next, an ontology is proposed that will allow for standardizing how data is transferred between the most common discipline-fidelity combinations used in the frameworks. A common representation of data creates the ability to 1) share system descriptions and analysis results, supporting more transparent benchmarks and comparison, and 2) integrate models together into workflows within and across organizations for improving the efficiency and performance of wind turbine and power plant design processes. Ultimately, this integration leads to better overall wind energy system designs with high performance and low costs.

17 WIND ENERGY↗

Development of a Novel DIW-PuSL Printer: The xPuSL Project

During my internship at Lawrence Livermore National Laboratory, I contributed to the xPuSL project. The xPuSL project aimed to revolutionize multi-material 3D printing by combining DIW and PuSL techniques to produce complex geometries with functional materials. My role encompassed process engineering, control systems, experimentation, and CAD design, leading to a streamlined workflow from CAD models to xPuSL prints.

36 MATERIALS SCIENCE↗

High-Fidelity Building Emulator for Integrated Comfort and Energy Analysis using EnergyPlus and Radiance

The growing need for smart, energy-efficient, and occupant-centric buildings has created a demand for advanced control systems that can optimize building operations to balance energy savings, demand flexibility, and comfort. However, current building energy simulation tools, such as EnergyPlus, have limitations that hinder the development and evaluation of these complex control systems. To address this challenge, we introduce a high-fidelity building emulator that dynamically couples EnergyPlus with Radiance for enhanced daylight modeling. The introduced workflow allows researchers and practitioners to rapidly develop and evaluate innovative control solutions. An example study looking at a south-facing office zone revealed up to 67% deviation in predicted light levels, which can significantly impact building assessment.

Yu, Tammie↗

United Space Alliance LLC Parachute Refurbishment Facility Model

The Parachute Refurbishment Facility Model was created to reflect the flow of hardware through the facility using anticipated start and delivery times from a project level IV schedule. Distributions for task times were built using historical build data for SFOC work and new data generated for CLV/ARES task times. The model currently processes 633 line items from 14 SFOC builds for flight readiness, 16 SFOC builds returning from flight for defoul, wash, and dry operations, 12 builds for CLV manufacturing operations, and 1 ARES 1X build. Modeling the planned workflow through the PRF is providing a reliable way to predict the capability of the facility as well as the manpower resource need. Creating a real world process allows for real world problems to be identified and potential workarounds to be implemented in a safe, simulated world before taking it to the next step, implementation in the real world.

Esser, Valerie↗

Modeling the Worst-Case-Scenario of Soil Drydown for Agricultural Drought Monitoring in Kansas Using NASA Earth Observations and In-situ Parameters

Climatologists in Kansas have observed rapid shifts from wet to dry conditions and rapid intensification of drought conditions for the state in recent years. As a leading state in agricultural production, drought events can lead to decreased crop yields and impact the state’s economy. The exponential decay of soil moisture content is a major consequence of drought and can exacerbate these impacts on agricultural production. This study examined the rate of soil moisture drydown in Kansas to understand and monitor the worst-case-scenario of short-term soil water shortage. Soil Moisture Active Passive (SMAP) L-band Radiometer rootzone soil moisture imagery from March 2015 to June 2019 was used to estimate the minimum and maximum soil moisture capacity throughout the state. In situ soil moisture parameters, such as the exponential decay constant, were approximated from the Kansas Mesonet system. These variables were inputs into an exponential decay model that forecasted gridded rootzone soil moisture in a scenario that assumed no precipitation for a user-identified period of time. The forecasts were output on the order of weeks to capture the rapid shifts in soil moisture content. Another model was used to identify areas in Kansas currently below a given percentage of the relative soil moisture saturation, measured as a function of minimum, maximum, and current soil moisture. This model forecasted the number of days until each cell in the soil moisture grid reached the input percentage of relative soil moisture saturation. The gridded model outputs and workflow were developed in partnership with the Kansas Water Office and Kansas Office of the State Climatologist at Kansas State University, as a part of their drought monitoring efforts.

NASA DEVELOP↗

Application of Agile for Systems Engineering, Project Management and Modeling and Lessons Learned

The Systems Engineering team within the Human Research Program (HRP) Exploration Medical Capability (ExMC) Element has been transforming its development processes to be more efficient, robust,and responsive to change and to its stakeholders. To these ends, the Systems Engineering team trialed the integration of agile development techniques into existing and new processes. Agile development methods are well understood within the software development community. Outside ofsoftware development, however, how non-software project management (PM) and systems engineering (SE) teams implement agile development techniques is less well understood. In its transformation efforts, the ExMC SE team focused on three main areas: Improving the project communications among subsystem teams and stakeholders by adopting a scrum-like process, Changing the status and reporting mechanisms to improve schedule coordination between the subsystem team, SE leadership, and ExMC Element leadership, and Unifying the model-based SE workflow to improve understanding of Concepts of Operations across projects.This presentation highlights several of these transformations and what the SE team learned while undergoing the transformation.

S Lumpkins↗

TEAL Output Visualization

Tool for Economic Analysis (TEAL) is an open-source economics calculation package written and maintained by Idaho National Laboratory (INL). As a plugin of the Risk Analysis Virtual ENvironment (RAVEN), TEAL provides economics analysis models to RAVEN workflow users. In addition, TEAL is used in other RAVEN plugins such as LOGOS and the Holistic Energy Resource Optimization Network (HERON) to complement other analyses with economic calculations. This report documents the efforts in developing the capability of TEAL output visualization for users to better understand and investigate TEAL simulation results. In the past, outputs (e.g., various cash flows) were generated and printed to screen. According to the need and purpose for visualization, users can select various bar charts and donut charts to visualize the cash flows including inflows, outflows, and net cash flows in each year or in a selected time range, in addition to observing data stored in comma-separated value files (CSVs). The visualization capability is flexible and user-friendly by introducing several user-defined parameters, for instance, the start and the end project year of interest and the type of colors to be assigned to the cash flows.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗