Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modeling workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Segmentation method comparison for residual fiber length measurement across tiled microscopy images

Fiber length distribution (FLD), in part, governs mechanical properties in discontinuous fiber composites, yet manual measurement methods limit the high-throughput characterization needed for materials design optimization. This study compares deep learning segmentation approaches for automated FLD measurement in large-field microscopy, evaluating how method choice affects the microstructural descriptors used in structure-property-processing relationships. A critical challenge is that high-resolution microscopy images (10,000×10,000 pixels) must be tiled for deep learning analysis, fragmenting fibers at boundaries. We demonstrate that segmentation method proves crucial for measurement accuracy. For example, instance segmentation with Slicing Aided Hyper Inference (SAHI) preserves individual fiber integrity across tiles while semantic segmentation prioritizes speed. Comparing against manual measurement of extracted carbon fibers, YOLOv11-SAHI matched manual ground truth (238 μm weighted mean) with 40x speedup (4.5 vs 167 minutes per image). U-Net provides rapid quantification although it is at the cost of reduced accuracy due only reliably measuring stand-alone fibers. Our comparative analysis reveals that instance segmentation with SAHI better preserves length measurements while semantic segmentation prioritizes speed, providing empirical guidance for method selection. The characterization provides essential inputs for mechanical property prediction models and inverse design workflows, accelerating composite materials development cycles.

Additive manufacturing↗

Computational Discovery of Intermolecular Singlet Fission Materials Using Many-Body Perturbation Theory

Intermolecular singlet fission (SF) is the conversion of a photogenerated singlet exciton into two triplet excitons residing on different molecules. SF has the potential to enhance the conversion efficiency of solar cells by harvesting two charge carriers from one high-energy photon, whose surplus energy would otherwise be lost to heat. The development of commercial SF-augmented modules is hindered by the limited selection of molecular crystals that exhibit intermolecular SF in the solid state. Computational exploration may accelerate the discovery of new SF materials. The GW approximation and Bethe–Salpeter equation (GW+BSE) within the framework of many-body perturbation theory is the current state-of-the-art method for calculating the excited-state properties of molecular crystals with periodic boundary conditions. In this Review, we discuss the usage of GW+BSE to assess candidate SF materials as well as its combination with low-cost physical or machine learned models in materials discovery workflows. We demonstrate three successful strategies for the discovery of new SF materials: (i) functionalization of known materials to tune their properties, (ii) finding potential polymorphs with improved crystal packing, and (iii) exploring new classes of materials. In addition, three new candidate SF materials are proposed here, which have not been published previously.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bridging Hydrological Ensemble Simulation and Learning Using Deep Neural Operators

Ensemble-based simulation and learning (ESnL) has long been used in hydrology for parameter inference, but computational demands of process-based ESnL can be quite high. To address this issue, we propose a deep neural operator learning approach. Neural operators are generic machine learning algorithms that can learn functional mappings between infinite-dimensional spaces, providing a highly flexible tool for scientific machine learning. Our approach is built upon DeepONet, a specific deep neural operator, and is designed to address several common problems in hydrology, namely, model parameter estimation, prediction at ungaged locations, and uncertainty quantification. Here we demonstrate the effectiveness of our DeepONet-based workflow using an existing large model ensemble created for an eastern U.S. watershed that is instrumented with 10 streamflow gages. Results suggest DeepONet achieves high efficiency in learning an ML surrogate model from the model ensemble, with the modified Kling-Gupta Efficiency exceeding 0.9 on holdout test sets. Parameter inference, carried out using the trained DeepONet surrogate model and genetic algorithm, also yields robust results. Additionally, we formulate and train a separate DeepONet model for physics-informed, seq-to-seq streamflow forecasting, which further reduces biases in the pre-trained DeepONet surrogate model. While this study focuses primarily on a single watershed, our approach is general and may be extended to enable learning from model ensembles across multiple basins or models. Thus, this research represents a significant contribution to the application of hybrid machine learning in hydrology.

54 ENVIRONMENTAL SCIENCES↗

Assessing the numerical stability of physics models to equilibrium variation through database comparisons on DIII-D

High fidelity kinetic equilibria are crucial for tokamak modeling and analysis. Manual workflows for constructing kinetic equilibria are time consuming and subject to user error, motivating development of automated equilibrium reconstruction tools to provide accurate and consistent reconstructions for downstream physics analysis. These automated tools also provide access to kinetic equilibria at large database scales, which enables the quantification of general uncertainties arising from equilibrium reconstruction techniques. In this paper, we compare a large database of DIII-D kinetic equilibria generated manually by physics experts to equilibria from automated kinetic reconstruction tools, assessing the impact of reconstruction method on equilibrium parameters and resulting magnetohydrodynamic stability calculations. We find agreement among scalar parameters, whereas profile quantities, such as the bootstrap current, show larger disagreements. We analyze ideal kink and classical tearing stability with DCON and STRIDE respectively, finding that the kink stability calculation is generally more robust than the tearing index Δ' calculation. We find that in 90% of cases, both kink stability classifications are unchanged between the manual expert and automated kinetic equilibria.

CAKE↗

Fast ion transport by sawtooth instability in the presence of ICRF-NBI synergy in JET plasmas

JET experiments have shown that the three-ion scenarios using waves in the ion cyclotron range of frequencies (ICRF) is an efficient way to build fast ion population through beam ion acceleration by radio frequency (RF) waves. Here, such a heating scheme is applied to plasmas with at least two thermal ion species. Analysis of mixed discharges with complex heating schemes requires a workflow that allows to model thermal and fast ion transport consistently. This paper is dedicated to modelling of a mixed plasma discharge with significant fraction of fast ions and contributes to development of fast ion transport models. For interpretive analysis with the TRANSP code a JET hydrogen-deuterium (HD) plasma discharge with neutral beam injection (NBI) and ICRF heating has been chosen. The task is complicated by NBI-ICRF synergy and plasma magnetohydrodynamic activity, like sawtooth crashes. D beam ions accelerated by RF waves form a high energy tail in fast ion distribution. Significant difference between the neutron rate computed by TRANSP and measured one is observed if the same diffusivity for electrons and ions is assumed. Sensitivity studies show that uncertainties in input plasma parameters and thermal ion transport models are crucial for modelling mixed plasma discharges and increased D transport is required to reach the plasma composition consistent with diagnostic measurements at the plasma edge. Fast ion redistribution by a sawtooth instability is characterised by non-resonant transport due to reconnection of magnetic field lines and resonant transport caused by resonance interaction between the instability and fast ions. With ORBIT simulations it has been shown that resonant interaction strongly affects fast ions of high energies, like beam ions accelerated by RF waves and fusion products. For the considered case, fast ion profiles simulated by ORBIT remain peaked after the sawtooth crashes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

LowFive v1.0

LowFive is a new data transport layer based on the HDF5 data model, for in situ workflows. Executables using LowFive can communicate in situ (using in-memory data and MPI message passing), reading and writing traditional HDF5 files to physical storage, and combining the two modes. Minimal and often no source-code modification is needed for programs that already use HDF5. LowFive maintains deep copies or shallow references of datasets, configurable by the user. More than one task can produce (write) data, and more than one task can consume (read) data, accommodating fan-in and fan-out in the workflow task graph. LowFive supports data redistribution from n producer processes to m consumer processes.

Morozov, Dmitriy↗

CDL2PLC translator v0.1.0

The CDL-PLC translator aims at translating control sequences for building energy systems from the CDL CXF format to the PLCopen XML format. The CDL CXF developed at LBL within the OpenBuildingControl project, and now being standardized via ASHRAE Standard 231P, enables expressing control sequences developed in the simulation environment Modelica in a JSON format. The PLCopen XML is an existing exchange format standardized in IEC 61131-10 for Programmable Logic Controllers (PLCs) following the IEC 61131 standard as one target system of CDL among others. The translation from the CDL CXF to the PLCopen XML contributes to a seamless workflow from the model-based development of control sequences in simulation environments, which is not building practice today, and their digital implementation on building controllers, which replaces graphical and textual documents used for this purpose today. The translator is at a prototypical stage and enables, as a proof of concept, the translation of very simple control sequences composed of 4 selected function blocks out of 137 function blocks defined in CDL. The translation includes the connection of inputs and outputs of function blocks and the expression of a control function in CDL to the equivalent code in IEC 61131-3.

Walther, Karl↗

CVEVOLVE

CVEvolve is an agentic AI system for autonomous algorithm discovery for scientific data processing. It creates workflows where large language model agents freely set up and configure development environments and evaluation harnesses, develop and improve data processing algorithms with designed exploration-exploitation balancing mechanisms, log history and findings in a structured database, and run holdout testing to ensure algorithm generalizability. CVEvolve offers a zero-code interface and does not require users to provide structured data and evaluation scripts.

Cherukara, MatthewJoseph [Argonne National Laborat↗

OpenOA: An Open-Source Codebase For Operational Analysis of Wind Farms

OpenOA is an open source framework for operational data analysis of wind energy plants, implemented in the Python programming language. OpenOA provides a common data model, high level analysis workflows, and low-level convenience functions that engineers, analysts, and researchers in the wind energy industry can use to facilitate analytics workflows on operational data sets. OpenOA contains documentation, worked out examples in Jupyter notebooks, and a corresponding example dataset from the Engie Renewable’s La Haute Borne Dataset.

17 WIND ENERGY↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Integrated Life Cycle and Techno-Economic Assessments of Central Appalachian Legacy Mine Sites for Biomass Development and Waste Coal Utilization

This project, funded by the U.S. Department of Energy – National Energy Technology Laboratory (DOE-NETL) under award DE-FE0032212, evaluated how legacy coal mine lands and coal refuse piles in Central Appalachia (West Virginia and Pennsylvania) can be reclaimed and repurposed to support biomass development and beneficial utilization of waste coal, with the long-term goal of supporting net-zero or net-negative greenhouse gas (GHG) pathways. The project had two primary objectives: 1. Characterize legacy mine sites (including site conditions, waste coal/refuse resources, and soil/ecosystem indicators) and develop reclamation and best management practices (BMPs) for biomass cultivation; and 2. Conduct integrated machine learning (ML)-assisted life cycle assessment (LCA) and techno-economic analysis (TEA) to quantify environmental and economic outcomes for multiple biomass and waste-coal utilization pathways. Across West Virginia, the team identified ~625 coal refuse sites covering ~19,705 acres, and developed methods to estimate refuse pile volume using digital elevation models (DEMs) and geospatial workflows. A large subset of sites received volume estimates totaling ~1.6 billion m³.

01 COAL, LIGNITE, AND PEAT↗

From Ensemble Climate to Ensemble Impacts

Many climate-risk tools rely on ensemble mean projections or endpoint climate snapshots to characterize future hazards. Although convenient for communication, these representations remove the statistical, temporal, and physical information that real infrastructure systems respond to. Infrastructure degradation and failure arise from extremes, sequences, cumulative stress, compound hazards, and nonlinear fragility relationships, none of which survive ensemble averaging or temporal compression. Power-system failure statistics and cascading failure models further show that infrastructure risk is dominated by tail events and path-dependent dynamics rather than by mean conditions. This paper demonstrates why ensemble mean or endpoint-only climate representations are mathematically and physically inconsistent with engineering-grade risk analysis. We outline a model-resolved, time-series-based workflow that preserves extremes, variability, and sequencing by propagating each climate-model realization independently through hazard formation, exposure, fragility, and cascading failure mechanisms. Taking the ensemble of impacts—rather than the ensemble of climate—provides a defensible, physically coherent foundation for infrastructure resilience planning, regulatory compliance, and long-term investment decisions.

54 - ENVIRONMENTAL SCIENCES/GLOBAL CLIMATE CHANGE ↗

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]↗

Development of Prototypical District-Scale Models

The U.S. has set the climate goal to achieve net-zero greenhouse gas emissions by 2050. District-scale solutions, which include scale-specific opportunities for energy and emissions savings, can be investigated and implemented to help accelerate decarbonization and progress toward this goal. However, there is currently a lack of district-scale models of buildings and community energy systems that can be used to evaluate potential district-scale technologies and strategies across a range of representative community types. This initial work aims to define and develop prototype district models that can be adapted to support the planning, design, and operation of buildings and energy systems in districts considering the complexity and interactions of diverse building loads, weather impacts, distributed energy resources (e.g., PV, EV, electric and thermal energy storage), electric and thermal grid systems, and pricing signals. An overall workflow for developing these prototype district models is established. Stakeholders and potential users of the prototype district models provided technical feedback. The specifications of the selected high priority districts were defined and documented in a scorecard format. An example prototype district model was implemented with the URBANopt platform workflows. A case study was performed to demonstrate the model application.

building energy modeling↗

Development of Prototypical District-Scale Models: Preprint

The U.S. has set the climate goal to achieve net-zero greenhouse gas emissions by 2050. District-scale solutions, which include scale-specific opportunities for energy and emissions savings, can be investigated and implemented to help accelerate decarbonization and progress toward this goal. However, there is currently a lack of district-scale models of buildings and community energy systems that can be used to evaluate potential district-scale technologies and strategies across a range of representative community types. This initial work aims to define and develop prototype district models that can be adapted to support the planning, design, and operation of buildings and energy systems in districts considering the complexity and interactions of diverse building loads, weather impacts, distributed energy resources (e.g., PV, EV, electric and thermal energy storage), electric and thermal grid systems, and pricing signals. An overall workflow for developing these prototype district models is established. Stakeholders and potential users of the prototype district models provided technical feedback. The specifications of the selected high priority districts were defined and documented in a scorecard format. An example prototype district model was implemented with the URBANoptTM platform workflows. A case study was performed to demonstrate the model application.

district↗

WATTS: Workflow and template toolkit for simulation

Modeling and simulation in many science and engineering domains often involves the execution and/or iteration of a sequence of applications, with data transfer between applications typically required. These applications often do not have a formal application programming interface (API). Instead, executing an application requires first writing a text-based input file, the format of which is typically defined in a user’s manual. While text-based input files are suitable for simple one-off calculations, they can become cumbersome if a user wants to execute the applications multiple times and systematically vary input parameters, especially when a complex workflow is involved. In this case, they must resort to either manually making changes in the input file or developing their own script that modifies the input file and executes the application. Depending on the format of the input file, writing such a script can be a non-trivial and error-prone task.

97 MATHEMATICS AND COMPUTING↗