Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Modeling workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Multiscale Machine-Learned Modeling Infrastructure

The Multiscale Machine-Learned Modeling Infrastructure (MuMMI) is a multiscale workflow management infrastructure that can concurrently orchestrate thousands of molecular dynamics (MD) simulations operating at different time and/or length scales, spanning nanoseconds to seconds and nanometers to micrometers. MuMMI uses machine learning (backed by biology experiments) to guide a massive ensemble of MD simulations that capture biologically relevant time and length scales with unprecedented resolution. MuMMI supports multiple MD codes such as GROMACS and ddcMD and can be fully deployed using the HPC package manager Spack. MuMMI has been used in many publications to run hundreds of thousands simulations, leading to significant biology breakthroughs.

Di Natale, Francesco [Lawrence Livermore National ↗

Navier: Dataflow Architecture for Computation Chemistry

Navier’s objectives were two evaluate the use of emerging technologies, especially dataflow accelerators, for high-performance computing (HPC) applications, specifically in the domain of chemistry, and to develop a prototype software stack to support such applications. Navier builds on capabilities previously developed by synergistic projects, such as PNNL Data Model Convergence (DMC) LDRD Hardware Advanced Workflows (HAW) and DuOMO, as well as DOE ARIAA. Throughout its 18 months, the Navier team developed new capabilities and artifacts at all levels of the HW/SW stack, provided a seamless way to integrate novel computing architectures (Sambanova SN10 and Xilinx Versal AI) into an existing software stack, developed chemistry workflows, data analytics tools, and HPC molecular dynamics workflows that leverage the developed stack and PNNL institutional investments in emerging architectures. Navier also explored the use of active learning to accelerate a computational chemistry workflow for organic molecules on PNNL Junction cluster (in collaboration with AMD/Xilinx). Navier developed tools, methodologies, and studies for hardware software co-design and (sparse) dataflow accelerators that are composable and can be used together or separately. These methodologies are now used in other projects, such as DOE AMAIS and HPDA. This report describes Navier’s achievement, the developed tools and methodologies, and the research findings and conclusions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

INTEGRATION OF DATA ANALYTICS WITH SYSTEM HEALTH PROGRAMS

Industry equipment reliability and asset management programs are essential elements that help ensure the safe and economical operation of nuclear power plants. The effectiveness of these programs is addressed in several industry developed and regulatory programs. However, these programs have proven to be labor intensive and expensive. There is an opportunity to significantly enhance the collection, analysis, and use of this information to provide more cost-effective plant operation. Additionally, there is an acute industry need to leverage advanced technology to reduce costs and improve operational effectiveness. The goal of this paper is to provide effective and efficient analytical methods and tools to support risk-informed decisions for the equipment reliability and asset management programs at nuclear power plants. This is accomplished by creating a direct bridge between component health/lifecycle data and decision making (e.g., maintenance scheduling and project prioritization). Here we are supporting typical system engineer decisions regarding maintenance activity scheduling and component ageing management. This is performed in a risk-informed context where herein the term “risk” is broadly constructed to include both plant reliability and economics. This framework combines data analytics tools to analyze equipment reliability data with risk-informed methods designed to support system engineer decisions (e.g., maintenance and replacement schedules, optimal maintenance posture) in a customizable workflow. A challenge is that the structure of this workflow strongly depends on the decision that needs to be made, the type of data available, and the constraints that need to be considered. Current methods are designed to provide specific answers to specific problems; however, these methods might prove to be inadequate even when problem settings slightly change (e.g., different types of requirements, additional dependencies between system reliability and economics). We tackled this challenge by designing framework in a flexible and modular fashion such that the user can assemble and customize his/her own workflow that integrates SSC economic lifecycle models (e.g., maintenance and replacement costs), system reliability models, and optimization methods.

97 - MATHEMATICS AND COMPUTING↗

Prediction of local concentration fields in porous media with chemical reaction using a multi scale convolutional neural network

The study of solute transport in porous media is of interest in many chemical engineering systems. Some example applications include packed bed catalytic reactors, filtration devices, and batteries. The pore scale modeling of these systems is time consuming and may require large computing resources, for this reason computational fluid dynamics (CFD) simulations are not practical if a large number of simulations is required, like in multiscale modeling, where a model at a large scale calls for pore scale simulations. It has been shown that neural networks can be trained with a dataset of flow simulations and then predict fields orders of magnitude faster, and with less computational resources, in new domains. However, it is crucial to provide the neural network with an effective description of the domain and the undergoing operating conditions to be able to train models that generalize accurately in unseen samples. Therefore, research is needed to employ neural networks in new complex systems. The appropriate training of a network for predicting coupled flow and solute transport processes is an outstanding problem due to the complex interplay between geometry and operating conditions. In this work, we train a multi scale convolutional neural network (MSNet) with a diverse dataset of simulations of transport and chemical reaction in porous media to predict the local concentration fields in images of porous media. Our dataset contains a wide diversity of sphere pack arrangements under different operating conditions (Péclet and Reynolds numbers). Further, we train a robust model by employing different input descriptors that represent the medium and the different operating conditions of each system. Our trained model is able to provide nearly instantaneous predictions, compared to around twenty hours of the CFD workflow, with less than 3.5% error on new geometries and transport conditions. Thus the model could be easily integrated in a multiscale workflow where fast response is needed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Description of FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment

The urgency to deliver fusion power is growing now more than ever, with increasing pressure for both public programs and private companies to meet milestones timelines and overcome significant remaining technical challenges to ensure growth of a nascent fusion industry in time to meet rapidly growing clean energy demands. With incredible advancements in computation and years of investment in fusion model development and validation, integrated modeling is poised to fill a key role in accelerating the timeline to a fusion pilot plant (FPP). Future fusion pilot plants will operate in regimes far beyond current experience, and device design will rely on physics-based prediction and extrapolation. Many concepts will also rely on simulation to assess safety (shielding, tritium management, materials activation and lifetimes), economics and scalability before the decision to build. Importantly, integrated simulation can be used to reveal and solve the complexities of system integration that may otherwise not be apparent in physical components or models developed in isolation. New experimental test facilities that produce relevant conditions to validate and resolve key technical challenges for various subsystems (materials, blankets, fuel cycle, etc.) have been repeatedly called for by the fusion community but are not yet realized. Integrated modeling has an important role in identifying realistic load conditions (thermal, electromagnetic, plasma, neutron and photon loads, etc.) and defining the components and experiments for these test facilities in order to ensure meaningful validation that sufficiently reduces modeling uncertainties and technical risk for the full integrated reactor. The Fusion REactor Design and Assessment (FREDA) SciDAC project is building a component-based integrated modeling framework & data structure to enable self-consistent, multi-fidelity, iterative optimization workflows for the fusion reactor design process. FREDA aims to shorten the time to viable designs by providing a set of flexible workflows to support the various stages of the design process using an integrated model hierarchy, ranging from the simple analytic descriptions to the highest fidelity, theory-based plasma and engineering modeling developed by the fusion and fission communities. These tools are expected to be needed for timely support of FPP design in the milestone program and in the FIRE collaboratives. The plasma simulation backbone of FREDA is IPS-FASTRAN with newly developed coupled Core-Edge Pedestal-SOL (CESOL) workflows, which is being extended to the far-SOL region up to the plasma facing components. FREDA incorporates the FERMI engineering modeling suite and will enable self-consistent evaluation of the thermal shields, limiters, blanket, magnets, and other surrounding structures with predictions of temperatures, erosion, dpa, activation, tritium generation and transport, creep, corrosion, material degradation, etc. Parametric generation of 3D CAD enables rapid iteration of component geometry in response to plasma and loading specifications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

Deep learning workflow for the inverse design of molecules with specific optoelectronic properties

The inverse design of novel molecules with a desirable optoelectronic property requires consideration of the vast chemical spaces associated with varying chemical composition and molecular size. First principles-based property predictions have become increasingly helpful for assisting the selection of promising candidate chemical species for subsequent experimental validation. However, a brute-force computational screening of the entire chemical space is decidedly impossible. To alleviate the computational burden and accelerate rational molecular design, we here present an iterative deep learning workflow that combines (i) the density-functional tight-binding method for dynamic generation of property training data, (ii) a graph convolutional neural network surrogate model for rapid and reliable predictions of chemical and physical properties, and (iii) a masked language model. As proof of principle, we employ our workflow in the iterative generation of novel molecules with a target energy gap between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO).

97 MATHEMATICS AND COMPUTING↗

A Generative Model for Synthetic Electroluminescence Images

This work will develop modular, open-source model and analysis components including crack detection workflow and parameterization for quantitative inspection of large EL large datasets. These tools will allow users to quickly and accurately assess the extent and types of cracking in their modules. Measured statistical distributions of crack parameters, together with the imposed stress and electrical properties will be used to generate models to predict future crack behavior and power loss.

Pierce, Benjamin Garrett↗

eXtremeMAT: Uncertainty Quantification of the LApx Model

Presentation on the uncertainty quantification efforts to parameterize and fit the LApx model for ferritic-martensitic steels and austenitic stainless steels. These workflows support ease of adoption of the code to new materials and enable rapid model fitting and understanding of the level of confidence in predictions.

mechanical properties↗

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

plexosdb: A Modular Library for Programmatic PLEXOS Model Construction

plexosdb is a lightweight Python library for constructing PLEXOS models using a SQLite-backed data structure. It provides a clear, modular interface that maps relational data directly to model components. By leveraging SQLite and idiomatic Python, it enables fast iteration and reproducible workflows. The result is a performant, composable foundation for scalable PLEXOS model development.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ChemGraph as an agentic framework for computational chemistry workflows

Atomistic simulations are essential in chemistry and materials science but remain challenging to run due to the expert knowledge required for the setup, execution, and validation stages of these calculations. We present ChemGraph, an agentic framework powered by artificial intelligence and state-of-the-art simulation tools to streamline and automate computational chemistry and materials science workflows. ChemGraph leverages graph neural network-based foundation models for accurate yet computationally efficient calculations and large language models (LLMs) for natural language understanding, task planning, and scientific reasoning to provide an intuitive and interactive interface. We evaluate ChemGraph across 13 benchmark tasks and demonstrate that smaller LLMs (GPT-4o-mini, Claude-3.5-haiku, Qwen-2.5-14B) perform well on simple workflows, while more complex tasks benefit from using larger models. Importantly, we show that decomposing complex tasks into smaller subtasks through a multi-agent framework enables GPT-4o to reach perfect accuracy and smaller LLMs to match or exceed single-agent GPT-4o's performance in these benchmarks.

Computational chemistry↗

Reactive transport modeling for supporting climate resilience at groundwater contamination sites

Abstract. Climate resilience is an emerging issue at contaminated sites and hazardous waste sites, since projected climate shifts (e.g., increased/decreased precipitation) and extreme events (e.g., flooding, drought) could affect ongoing remediation or closure strategies. In this study, we develop a reactive transport model (Amanzi) for radionuclides (uranium, tritium, and others) and evaluate how different scenarios under climate change will influence the contaminant plume conditions and groundwater well concentrations. We demonstrate our approach using a two-dimensional (2D) reactive transport model for the Savannah River Site F-Area, including mineral reaction and sorption processes. Different recharge scenarios are considered by perturbing the infiltration rate from the base case as well as considering cap-failure and climate projection scenarios. We also evaluate the uranium and nitrate concentration ratios between scenarios and the base case to isolate the sorption effects with changing recharge rates. The modeling results indicate that the competing effects of dilution and remobilization significantly influence pH, thus changing the sorption of uranium. At the maximum concentration on the breakthrough curve, higher aqueous uranium concentration implies that sorption is reduced with lower pH due to remobilization. To better evaluate the climate change impacts in the future, we develop the workflow to include the downscaled CMIP5 (Coupled Model Intercomparison Project) climate projection data in the reactive transport model and evaluate how residual contamination evolves through 2100 under four climate Representative Concentration Pathway (RCP) scenarios. The integration of climate modeling data and hydrogeochemistry models enables us to quantify the climate change impacts, assess which impacts need to be planned for, and therefore assist climate resiliency efforts and help guide site management.

54 ENVIRONMENTAL SCIENCES↗

CASTELO: clustered atom subtypes aided lead optimization—a combined machine learning and molecular modeling method

Background: Drug discovery is a multi-stage process that comprises two costly major steps: pre-clinical research and clinical trials. Among its stages, lead optimization easily consumes more than half of the pre-clinical budget. We propose a combined machine learning and molecular modeling approach that partially automates lead optimization workflow in silico, providing suggestions for modification hot spots. Results: The initial data collection is achieved with physics-based molecular dynamics simulation. Contact matrices are calculated as the preliminary features extracted from the simulations. To take advantage of the temporal information from the simulations, we enhanced contact matrices data with temporal dynamism representation, which are then modeled with unsupervised convolutional variational autoencoder (CVAE). Finally, conventional and CVAE-based clustering methods are compared with metrics to rank the submolecular structures and propose potential candidates for lead optimization. Conclusion: With no need for extensive structure-activity data, our method provides new hints for drug modification hotspots which can be used to improve drug potency and reduce the lead optimization time. It can potentially become a valuable tool for medicinal chemists.

59 BASIC BIOLOGICAL SCIENCES↗

Complete Demonstration of a Prototype Version of FORCE User Interface and Conduct Analyst Survey Collecting Feedback on Interface Features and Usability

In 2024 the US Department of Energy (DOE) Office of Nuclear Energy (NE) Integrated Energy System (IES) program continued to develop the Framework for Optimization of Resources and Economics (FORCE) analysis ecosystem into a more traditional toolset with simplified software installation, automated workflows, and interactive results visualization. The DOE-NE Nuclear Energy Advanced Modeling and Simulation (NEAMS) Workbench continued to be leveraged for user input, application workflow and runtime environment, and interactive results visualization capabilities. This report documents the demonstration of a FORCE User Interface (UI) prototype and the results of a survey of analysts’ using the Holistic Energy Resource Optimization Network (HERON) tool in FORCE with the prototype UI.

97 MATHEMATICS AND COMPUTING↗

Investigation of CAD-based Geometry Workflows for Multiphysics Fusion Problems Using OpenMC and MOOSE

Fusion system designs are complex and require intricate and accurate meshes to be properly modeled. In this study, we investigate the use of CAD-based geometry workflows in fusion systems multiphysics problems. A simplified tokamak was introduced and modeled in CAD using a multiphysics coupling of OpenMC Monte Carlo transport and MOOSE heat conduction. The meshed geometry was prepared using direct accelerated geometry Monte Carlo (DAGMC) for particle transport, and a volumetric mesh was also prepared to be used in MOOSE and to tally OpenMC results. Cardinal was used to run OpenMC Monte Carlo particle transport within MOOSE framework. The heat source distribution and tritium production were calculated in OpenMC. The data transfer system was used to transfer heat source and temperature distribution between OpenMC and MOOSE. Two computational studies related to mesh refinement were performed: (1) refining the DAGMC and volumetric meshes used for tallying results and solving heat conduction and (2) only refining the DAGMC particle transport mesh. The refinement of the tally mesh has a much larger effect on the runtime compared to the refinement of the DAGMC particle transport surface mesh.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Portable Programming Model Exploration for LArTPC Simulation in a Heterogeneous Computing Environment: OpenMP vs. SYCL

The evolution of the computing landscape has resulted in the proliferation of diverse hardware architectures, with different flavors of GPUs and other compute accelerators becoming more widely available. To facilitate the efficient use of these architectures in a heterogeneous computing environment, several programming models are available to enable portability and performance across different computing systems, such as Kokkos, SYCL, OpenMP and others. As part of the High Energy Physics Center for Computational Excellence (HEP-CCE) project, we investigate if and how these different programming models may be suitable for experimental HEP workflows through a few representative use cases. One of such use cases is the Liquid Argon Time Projection Chamber (LArTPC) simulation which is essential for LArTPC detector design, validation and data analysis. Following up on our previous investigations of using Kokkos to port LArTPC simulation in the Wire-Cell Toolkit (WCT) to GPUs, we have explored OpenMP and SYCL as potential portable programming models for WCT, with the goal to make diverse computing resources accessible to the LArTPC simulations. In this work, we describe how we utilize relevant features of OpenMP and SYCL for the LArTPC simulation module in WCT. We also show performance benchmark results on multi-core CPUs, NVIDIA and AMD GPUs for both the OpenMP and the SYCL implementations. Comparisons with different compilers will also be given where appropriate.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗