Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Powered By ERAD [Slides]

Energy Resilience Analysis for Distribution Power System (ERAD) is a free, open-source Python toolkit for estimating the energy and service impacts of hazards like earthquakes and flooding. It uses a graph-based approach to capture high resolution connectivity among the grid, critical services, and customers and rapidly compute household level metrics and aggregated statistics across large distribution systems. It uses asset fragility curves that relate hazard severity to survival probability for power system equipment including cables, transformers, substations, etc. The tool is designed to be modular and extensible, allowing it to interface with third-party hazard simulators and integrate into broader resilience analysis workflows. ERAD enables researchers, students, communities, distribution utilities, and other stakeholders to understand hazard impacts and evaluate the effectiveness of different programs to improve energy resilience. The webinar was hosted by NLR researcher Aadil Latif.

24 POWER TRANSMISSION AND DISTRIBUTION↗

DEIMoS: An Open-Source Tool for Processing High-Dimensional Mass Spectrometry Data

We present DEIMoS: Data Extraction for Integrated Multidimensional Spectrometry, a Python application programming interface and command-line tool for high-dimensional mass spectrometry data analysis workflows, offering ease of development and access to efficient algorithmic implementations. Functionality includes feature detection, feature alignment, collision cross section calibration, isotope detection, and MS/MS spectral deconvolution, with the output comprising detected features aligned across study samples and characterized by mass, CCS, tandem mass spectra, and isotopic signature. Notably, DEIMoS operates on N-dimensional data, largely agnostic to acquisition instrumentation: algorithm implementations utilize all dimensions simultaneously to (i) offer greater separation between features, improving detection sensitivity, (ii) increase alignment/feature matching confidence among datasets, and (iii) mitigate convolution artifacts in tandem mass spectra. We demonstrate DEIMoS with LC-IMS-MS/MS data, demonstrating the advantages of a multidimensional approach in each data processing step.

Colby, Sean M↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

Wavelet Analysis of GPR Data for Belowground Mass Assessment of Sorghum Hybrid for Soil Carbon Sequestration

Among many agricultural practices proposed to cut carbon emissions in the next 30 years is the deposition of carbon in soils as plant matter. Adding rooting traits as part of a sequestration strategy would result in significantly increased carbon sequestration. Integrating these traits into production agriculture requires a belowground phenotyping method compatible with high-throughput breeding (i.e., rapid, inexpensive, reliable, and non-destructive). However, methods that fulfill these criteria currently do not exist. We hypothesized that ground-penetrating radar (GPR) could fill this need as a phenotypic selection tool. In this study, we employed a prototype GPR antenna array to scan and discriminate the root and rhizome mass of the perennial sorghum hybrid PSH09TX15. B-scan level time/discrete frequency analyses using continuous wavelet transform were utilized to extract features of interest that could be correlated to the biomass of the subsurface roots and rhizome. Time frequency analysis yielded strong correlations between radar features and belowground biomass (max R −0.91 for roots and −0.78 rhizomes, respectively) These results demonstrate that continued refinement of GPR data analysis workflows should yield an applicable phenotyping tool for breeding efforts in contexts where selection is otherwise impractical.

Wolfe, Matthew↗

Optimisation of the Search for CP-symmetry Violation at the Deep Underground Neutrino Experiment

The Deep Underground Neutrino Experiment (DUNE) is a next-generation long baseline experiment, which will be situated in South Dakota. Its detectors will utilise liquid-argon time projection chamber technology, which is able to capture neutrino interactions with an incredible spatial and calorimetric resolution. With what will become the world’s most intense neutrino beam, a highly capable near detector, and four (10kt fducial mass) far detector modules, DUNE will be able to achieve an ambitious physics programme. Most notably, DUNE will determine whether charge-parity symmetry is broken in neutrino oscillations - a finding that would have significant implications for the understanding of the matter-antimatter asymmetry in our Universe.This thesis presents the optimisation of a CP-violation analysis at DUNE using thePandora pattern-recognition software. The analysis assumes a 3.5 year exposure (1.36 × 1023 protons on target) to a neutrino and an antineutrino beam (7 year total). Only the predicted data of the far detector modules is used; near detector samples are not included. The initial sensitivity to CP-violation is found to be 3.8σ+0.9σ−1.1σ in an estimate that includes oscillation parameter uncertainties, systematic uncertainties and statistical fluctuations, and assumes a normal-ordering of the neutrino mass hierarchy. The performance of the Pandora event reconstruction is linked to that of the analysis, which is found to be limited by the reconstruction of the initial track-like region of electrons and photons. The ShowerRefinement algorithm is developed in response to this and its implementation into the analysis workflow results in an improved sensitivity to CP-violation of 4.6σ+0.9σ−1.0σ. With a perfected neutrino interaction vertex placement, this is further increased to 5.1σ+1.0σ−1.1σ.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Physics demonstration and verification of MOOSE framework reactor module meshing capabilities

The recently developed Reactor module in the open-source MOOSE framework includes finite element meshing capabilities for analysis of common reactor geometries. Capabilities in the Reactor module have been employed to generate meshes for physics applications including a sodium-cooled fast reactor core analysis using Griffin, a fast reactor assembly thermal deformation analysis using MOOSE Tensor Mechanics, and a heat-pipe cooled microreactor coupled analysis using Griffin, Bison, and Sockeye. The process to build these meshes using MOOSE's meshing capabilities is described. Physics simulation results using MOOSE-based meshes have been verified to match results which leverage external meshing software such as Cubit or Argonne's Mesh Tools. MOOSE's Reactor module provides significant advantages compared to the use of external meshing tools when analyzing Cartesian and hexagonal reactor lattices using MOOSE-based applications: accessibility to the end user, low barrier to entry for new users, speed of mesh generation, volume preservation of meshed fuel pins, and simplification of analysis workflow when using MOOSE applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A versatile machine learning workflow for high-throughput analysis of supported metal catalyst particles

Accurate and efficient characterization of nanoparticles (NPs), particularly regarding particle size distribution, is essential for advancing our understanding of their structure-property relationship and facilitating their design for various applications. In this study, we introduce a novel two-stage artificial intelligence (AI)-driven workflow for NP analysis that leverages prompt engineering techniques from state-of-the-art single-stage object detection and large-scale vision transformer (ViT) architectures. This methodology is applied to transmission electron microscopy (TEM) and scanning TEM (STEM) images of heterogeneous catalysts, enabling high-resolution, high-throughput analysis of particle size distributions for supported metal catalyst NPs. The model's performance in detecting and segmenting NPs is validated across diverse heterogeneous catalyst systems, including various metals (Ru, Cu, PtCo, and Pt), supports (silica (SiO 2 ), γ-alumina (γ-Al 2 O 3 ), and carbon black), and particle diameter size distributions with mean and standard deviations ranging from 1.6 ± 0.2 nm to 9.7 ± 4.6 nm. The proposed machine learning (ML) methodology achieved an average F1 overlap score of 0.91 ± 0.01 and demonstrated the ability to disentangle overlapping NPs anchored on catalytic support materials. The segmentation accuracy is further validated using the Hausdorff distance and robust Hausdorff distance metrics, with the 90th percent of the robust Hausdorff distance showing errors within 0.4 ± 0.1 nm to 1.4 ± 0.6 nm. In conclusion, our AI-assisted NP analysis workflow demonstrates robust generalization across diverse datasets and can be readily applied to similar NP segmentation tasks without requiring costly model retraining.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

Meshing, Language Server Protocol, and other User-Oriented MOOSE Framework Improvements to Enhance Reactor Analysis Capabilities and Workflows

The open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) framework underpins most of the Nuclear Energy Advanced Modeling and Simulation (NEAMS) physics applications and coupling methods. Users interact with MOOSE in multiple ways to enable the construction of complex multiphysics models, including but not limited to compilation, input creation and syntax validation, meshing, solving the physics problem via MOOSE-based solvers and coupling, and output inspection. In FY23, numerous enhancements have been made to MOOSE to enhance the user experience. Reactor-oriented meshing capabilities in MOOSE have been expanded, an online tutorial was created and hosted on the MOOSE site, and a hands-on workshop was held for more than 100 users. The Language Server Protocol (LSP) capability has been implemented in MOOSE to better communicate correct syntax to the NEAMS Workbench user interface which is commonly used to validate input and submit jobs by users. Finally, several other user-facing improvements were made including the implementation of improved screen output and logging control, enhancements to the initial condition system, and enhancements of the MultiApp system commonly used for coupling. Collectively, these enhancements were driven by user needs and directly improve user ability to construct complex multiphysics models for nuclear reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

From minimum-viable-products to full models: a step-wise development of diagnostic forward models in support of design, analysis and modelling on the ST40 tokamak

Like most magnetic confined fusion experiments, the ST40 tokamak started off with a small subset of diagnostics and gradually increased the diagnostic set to include more complex and comprehensive systems. To make the most of each operational phase, forward models of various diagnostics are used and developed to aid design, provide consistency-checks during commissioning, test analysis methods, and build workflows to constrain high-level parameters to inform interpretation, theory and modelling. For new models and new analysis workflows, minimum-viable-products are released early, and their complexity is increased in a step-wise manner, facilitating the support of all programme phases on multiple parallel applications, while enabling learning opportunities and feedback loops. In this contribution we review the philosophy, scope and architecture of the framework under development. We discuss the details of some forward models, with examples on how they are used to aid diagnostic design, to investigate analysis methodologies through synthetic data, and how they are embedded in experimental analysis workflows. We compare previously published experimental results with new, more advanced analysis workflows employing more recent, detailed models and new diagnostic data, providing confirmation of the published material from the 2021–22 experimental campaign.

integrated data analysis↗

FIRM image analysis: A machine learning workflow for quantifying extracellular matrix components from electron microscopy images

The extracellular matrix (ECM) is a complex network of biomolecules that plays an integral role in the structure, processes, and signaling mechanisms of cells and tissues. Identifying and quantifying changes in these matrix components provides insight into the mechanisms behind specific tissue remodeling processes; however, quantifying these changes is challenging due to difficult imaging conditions, complexity of the ECM, and the subtlety of these changes. Current imaging techniques allow us to visualize these critical remodeling events and developments in image analysis have employed a combination of analysis software and machine learning techniques to improve the efficiency and accuracy with which features are measured. Although image analysis has seen much improvement in recent years, there has been no technique developed to address ambiguity in feature edges in electron microscopy images. Presented here is a new machine learning-based workflow for the analysis of microscopy images named FIRM (Feature Identification from Raw Microscopy) that uses a random forest classifier to identify ECM features of interest and generate binary segmentation masks for quantification with ImageJ-FIJI. FIRM performed with an F1 score of 0.794 and greater than 80% accuracy for number and size of features detected. FIRM had similar deviation from the ground truth in the number of identified fibrils, fibril size, and size distributions when compared to human analyses. The results suggest that FIRM performs as well as manual analysis and requires a fraction of the time. This analysis technique is more efficient, eliminates user bias, and can be easily optimized to identify a variety of features, making it useful for any discipline requiring image analysis.

Science & Technology - Other Topics↗

Enhancing Monte Carlo Workflows for Nuclear Reactor Analysis with Metamodel-Driven Modeling

Monte Carlo codes are essential components of many reactor physics simulation workflows as high-fidelity continuous-energy neutron transport solvers. Among Monte Carlo radiation transport codes, MCNP is particularly notable due to its diverse simulation capabilities, large user base, and long validation history. Despite being a powerful simulation tool, MCNP provides limited capabilities to allow automated execution, model transformation, or support for user-defined logic and abstractions that limit its compatibility with modern workflows. Here, to better integrate MCNP into a modern scientific workflow, we have developed an intuitive yet full-featured MCNP Application Program Interface (API) in Python, named MCNPy, which provides a specialized set of classes for MCNP input development. Moreover, to guarantee that our reading, writing, and modeling capabilities remain self-consistent (and to render the huge scope of the MCNP API manageable), we have adopted a strategy of model-driven software development in which a generalized model of the MCNP input format has been created. From this generalized model, or “metamodel,” problem-specific implementations such as an engine for input validation or a codebase for programmatic operations may be automatically generated. Since MCNPy primarily acts as a Python front-end to the underlying Java API that directly interfaces with the metamodel, it is intrinsically linked to the metamodel and thus remains maintainable. With MCNPy, users can programmatically read, write, and modify any syntactically valid MCNP input file regardless of its origin. These capabilities allow users to automate complicated tasks like design optimization and model translation for nuclear systems. As examples, this work demonstrates the use of MCNPy to find the critical radius of a plutonium sphere and to translate a 9000+ line MCNP input file into a corresponding OpenMC model.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Workflow for High-throughput Screening of Enzyme Mutant Libraries Using Matrix-assisted Laser Desorption/Ionization Mass Spectrometry Analysis of Escherichia coli Colonies

High-throughput molecular screening of microbial colonies and DNA libraries are critical procedures that enable applications such as directed evolution, functional genomics, microbial identification, and creation of engineered microbial strains to produce high-value molecules. A promising chemical screening approach is the measurement of products directly from microbial colonies via optically guided matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). Measuring the compounds from microbial colonies bypasses liquid culture with a screen that takes approximately 5 s per sample. We describe a protocol combining a dedicated informatics pipeline and sample preparation method that can prepare up to 3,000 colonies in under 3 h. The screening protocol starts from colonies grown on Petri dishes and then transferred onto MALDI plates via imprinting. The target plate with the colonies is imaged by a flatbed scanner and the colonies are located via custom software. The target plate is coated with MALDI matrix, MALDI-MS analyzes the colony locations, and data analysis enables the determination of colonies with the desired biochemical properties. This workflow screens thousands of colonies per day without requiring additional automation. The wide chemical coverage and the high sensitivity of MALDI-MS enable diverse screening projects such as modifying enzymes and functional genomics surveys of gene activation/inhibition libraries.

Choe, Kisurb↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

ATLAS Data Analysis using a Parallel Workflow on Distributed Cloud-based Services with GPUs

A new type of parallel workflow is developed for the ATLAS experiment at the Large Hadron Collider, that makes use of distributed computing combined with a cloud-based infrastructure. This has been developed for a specific type of analysis using ATLAS data, one popularly referred to as Simulation-Based Inference (SBI). The JAX library is used for the parts of the workflow to compute gradients as well as accelerate program execution using just-in-time compilation, which becomes essential in a full SBI analysis and can also offer significant speed-ups in more traditional types of analysis.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enabling discovery data science through cross-facility workflows

Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.

Antypas, Katerina B.↗

Enhancement of PyARC for Westinghouse Electric Company’s Lead Fast Reactor Design and Modeling (Final TCF Report)

Westinghouse Electric Company is a nuclear reactor vendor headquartered in the U.S. that is developing advanced reactor technology for the U.S. and global markets. Westinghouse has been relying on the neutronics Argonne Reactor Codes (ARC) executed through the NEAMS Workbench and its PyARC module that are developed under the DOE-NE Nuclear Energy Advanced Modeling and Simulation (NEAMS) and Advanced Reactor Technology (ART) – Fast Reactor programs. Through this user experience, Westinghouse identified several enhancements that would benefit the ARC codes’ usability by the US industry and therefore its commercialization potential. The enhancements were proposed to deliver both improvements in workflow and analysis capabilities to better support effective fast reactor core design and analysis to the nuclear industry. The PyARC workflow was extended in this project by integrating non-neutronic ARC codes DASSH and NUBOW-3D. The Ducted Assembly Steady-State Heat equation (DASSH) code is developed at ANL to perform steady-state thermal hydraulic sub-channel analysis in liquid metal fast reactor assemblies to determine optimized coolant flow and temperature distributions, which in this project was updated and validated for lead fast reactor (LFR) applications. The interface between REBUS and NUBOW-3D were improved in this project to assess the impact of the core restraint design and thermal induced expansion effects on the reactivity of the core, and to model the deformations of the fuel assemblies induced by temperature and irradiation. Finally, the ARC models that were extensively verified and validated through various SFR-based modeling benchmarks are extended in this project through code-to-code comparison on relevant LFR-specific neutronics benchmarks against Monte-Carlo neutronic solutions. Overall, this work enables verification of the capability of the ARC codes for a wide range of Generation-IV reactor designs. The outcome of this project is the release of a comprehensive modeling toolkit of validated, robust and efficient codes, as well as their user interface, that enables industry to perform a wide range of fast reactor analyses for design and licensing of their concepts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗