Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Visual HPC Workflows for the Analysis of System Dynamics Models

Visual analytics supported by high performance computing (HPC) accelerates and enhances the discovery, exploration, and analysis of causal patterns in complex system dynamics (SD) models. We present a suite of visualization-assisted ensemble-based techniques for hypothesis generation and testing, and for sensitivity analysis. By employing HPC to provide parallel, on-demand simulation of SD models, one can “steer” an ensemble of simulated scenarios in real time as one first formulates and then informally tests those hypotheses: this provides rapid feedback for analysts to refine their understanding of the causal relationships emergent from a model. Such understandings can be followed and augmented by rigorous application of statistical methods, namely global variance-based sensitivity analysis, Monte-Carlo filtering, adaptive regional sensitivity analysis, and self-organized maps: here timely computation relies on HPC, while effective presentation emphasizes high-dimensional multivariate data visualization. Immersive visualization in virtual 3D environments provides an excellent adjunct to the traditional 2D graphics typically used for SD models, as it generates an embodied understanding of model behavior and facilitates an active, collaborative critique of model structure and output. Finally, we summarize prospects for HPC-enabled visual analytics applied to SD modeling.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Place polarized neutron reflectivity data in bins of constant wavevector transfer

Code to place polarized neutron reflectivity into bins of constant Qz with regard for Zeeman energy change to the spin-flip channel. The code is written as a Jupyter notebook. Example data are included as separate files (spin_down.txt and spin_up.txt). As a first step, run all cells in the notebook. Opportunities are identified for users to make their own choices for analysis of the data. The code is written to be a stand-alone operation on data collected from a pulse neutron source that have been previously normalized to the wavelength dependence of the incident beam spectrum, background subtracted, etc. The constant Qz binning (this code) is the last step to perform in the analysis (reduction) workflow. This step yields the the reflectivities for the four neutron spin cross-sections. A future intention is to create a version of this code that will be one module of many in the data analysis workflow.

Fitzsimmons, Michael [Oak Ridge National Lab. (ORN↗

From minimum-viable-products to full models: a step-wise development of diagnostic forward models in support of design, analysis and modelling on the ST40 tokamak

Like most magnetic confined fusion experiments, the ST40 tokamak started off with a small subset of diagnostics and gradually increased the diagnostic set to include more complex and comprehensive systems. To make the most of each operational phase, forward models of various diagnostics are used and developed to aid design, provide consistency-checks during commissioning, test analysis methods, and build workflows to constrain high-level parameters to inform interpretation, theory and modelling. For new models and new analysis workflows, minimum-viable-products are released early, and their complexity is increased in a step-wise manner, facilitating the support of all programme phases on multiple parallel applications, while enabling learning opportunities and feedback loops. In this contribution we review the philosophy, scope and architecture of the framework under development. We discuss the details of some forward models, with examples on how they are used to aid diagnostic design, to investigate analysis methodologies through synthetic data, and how they are embedded in experimental analysis workflows. We compare previously published experimental results with new, more advanced analysis workflows employing more recent, detailed models and new diagnostic data, providing confirmation of the published material from the 2021–22 experimental campaign.

integrated data analysis↗

FIRM image analysis: A machine learning workflow for quantifying extracellular matrix components from electron microscopy images

The extracellular matrix (ECM) is a complex network of biomolecules that plays an integral role in the structure, processes, and signaling mechanisms of cells and tissues. Identifying and quantifying changes in these matrix components provides insight into the mechanisms behind specific tissue remodeling processes; however, quantifying these changes is challenging due to difficult imaging conditions, complexity of the ECM, and the subtlety of these changes. Current imaging techniques allow us to visualize these critical remodeling events and developments in image analysis have employed a combination of analysis software and machine learning techniques to improve the efficiency and accuracy with which features are measured. Although image analysis has seen much improvement in recent years, there has been no technique developed to address ambiguity in feature edges in electron microscopy images. Presented here is a new machine learning-based workflow for the analysis of microscopy images named FIRM (Feature Identification from Raw Microscopy) that uses a random forest classifier to identify ECM features of interest and generate binary segmentation masks for quantification with ImageJ-FIJI. FIRM performed with an F1 score of 0.794 and greater than 80% accuracy for number and size of features detected. FIRM had similar deviation from the ground truth in the number of identified fibrils, fibril size, and size distributions when compared to human analyses. The results suggest that FIRM performs as well as manual analysis and requires a fraction of the time. This analysis technique is more efficient, eliminates user bias, and can be easily optimized to identify a variety of features, making it useful for any discipline requiring image analysis.

Science & Technology - Other Topics↗

Enhancing Monte Carlo Workflows for Nuclear Reactor Analysis with Metamodel-Driven Modeling

Monte Carlo codes are essential components of many reactor physics simulation workflows as high-fidelity continuous-energy neutron transport solvers. Among Monte Carlo radiation transport codes, MCNP is particularly notable due to its diverse simulation capabilities, large user base, and long validation history. Despite being a powerful simulation tool, MCNP provides limited capabilities to allow automated execution, model transformation, or support for user-defined logic and abstractions that limit its compatibility with modern workflows. Here, to better integrate MCNP into a modern scientific workflow, we have developed an intuitive yet full-featured MCNP Application Program Interface (API) in Python, named MCNPy, which provides a specialized set of classes for MCNP input development. Moreover, to guarantee that our reading, writing, and modeling capabilities remain self-consistent (and to render the huge scope of the MCNP API manageable), we have adopted a strategy of model-driven software development in which a generalized model of the MCNP input format has been created. From this generalized model, or “metamodel,” problem-specific implementations such as an engine for input validation or a codebase for programmatic operations may be automatically generated. Since MCNPy primarily acts as a Python front-end to the underlying Java API that directly interfaces with the metamodel, it is intrinsically linked to the metamodel and thus remains maintainable. With MCNPy, users can programmatically read, write, and modify any syntactically valid MCNP input file regardless of its origin. These capabilities allow users to automate complicated tasks like design optimization and model translation for nuclear systems. As examples, this work demonstrates the use of MCNPy to find the critical radius of a plutonium sphere and to translate a 9000+ line MCNP input file into a corresponding OpenMC model.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Workflow for High-throughput Screening of Enzyme Mutant Libraries Using Matrix-assisted Laser Desorption/Ionization Mass Spectrometry Analysis of Escherichia coli Colonies

High-throughput molecular screening of microbial colonies and DNA libraries are critical procedures that enable applications such as directed evolution, functional genomics, microbial identification, and creation of engineered microbial strains to produce high-value molecules. A promising chemical screening approach is the measurement of products directly from microbial colonies via optically guided matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). Measuring the compounds from microbial colonies bypasses liquid culture with a screen that takes approximately 5 s per sample. We describe a protocol combining a dedicated informatics pipeline and sample preparation method that can prepare up to 3,000 colonies in under 3 h. The screening protocol starts from colonies grown on Petri dishes and then transferred onto MALDI plates via imprinting. The target plate with the colonies is imaged by a flatbed scanner and the colonies are located via custom software. The target plate is coated with MALDI matrix, MALDI-MS analyzes the colony locations, and data analysis enables the determination of colonies with the desired biochemical properties. This workflow screens thousands of colonies per day without requiring additional automation. The wide chemical coverage and the high sensitivity of MALDI-MS enable diverse screening projects such as modifying enzymes and functional genomics surveys of gene activation/inhibition libraries.

Choe, Kisurb↗

Performance analysis and data reduction for exascale scientific workflows

Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.

97 MATHEMATICS AND COMPUTING↗

ATLAS Data Analysis using a Parallel Workflow on Distributed Cloud-based Services with GPUs

A new type of parallel workflow is developed for the ATLAS experiment at the Large Hadron Collider, that makes use of distributed computing combined with a cloud-based infrastructure. This has been developed for a specific type of analysis using ATLAS data, one popularly referred to as Simulation-Based Inference (SBI). The JAX library is used for the parts of the workflow to compute gradients as well as accelerate program execution using just-in-time compilation, which becomes essential in a full SBI analysis and can also offer significant speed-ups in more traditional types of analysis.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enabling discovery data science through cross-facility workflows

Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.

Antypas, Katerina B.↗

Enhancement of PyARC for Westinghouse Electric Company’s Lead Fast Reactor Design and Modeling (Final TCF Report)

Westinghouse Electric Company is a nuclear reactor vendor headquartered in the U.S. that is developing advanced reactor technology for the U.S. and global markets. Westinghouse has been relying on the neutronics Argonne Reactor Codes (ARC) executed through the NEAMS Workbench and its PyARC module that are developed under the DOE-NE Nuclear Energy Advanced Modeling and Simulation (NEAMS) and Advanced Reactor Technology (ART) – Fast Reactor programs. Through this user experience, Westinghouse identified several enhancements that would benefit the ARC codes’ usability by the US industry and therefore its commercialization potential. The enhancements were proposed to deliver both improvements in workflow and analysis capabilities to better support effective fast reactor core design and analysis to the nuclear industry. The PyARC workflow was extended in this project by integrating non-neutronic ARC codes DASSH and NUBOW-3D. The Ducted Assembly Steady-State Heat equation (DASSH) code is developed at ANL to perform steady-state thermal hydraulic sub-channel analysis in liquid metal fast reactor assemblies to determine optimized coolant flow and temperature distributions, which in this project was updated and validated for lead fast reactor (LFR) applications. The interface between REBUS and NUBOW-3D were improved in this project to assess the impact of the core restraint design and thermal induced expansion effects on the reactivity of the core, and to model the deformations of the fuel assemblies induced by temperature and irradiation. Finally, the ARC models that were extensively verified and validated through various SFR-based modeling benchmarks are extended in this project through code-to-code comparison on relevant LFR-specific neutronics benchmarks against Monte-Carlo neutronic solutions. Overall, this work enables verification of the capability of the ARC codes for a wide range of Generation-IV reactor designs. The outcome of this project is the release of a comprehensive modeling toolkit of validated, robust and efficient codes, as well as their user interface, that enables industry to perform a wide range of fast reactor analyses for design and licensing of their concepts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Parallel Machine Learning Workflow for Neutron Scattering Data Analysis

As part of a larger effort, this work-in-progress reports the possible advantages of modifying conventional workflows used to generate labelled training samples and train machine learning (ML) models on them. We compare results from three different workflows using neutron scattering data analysis as the motivating application and report about 20% improvement in speedup, with no appreciable loss of model accuracy, over a baseline workflow.

Wang, Tianle↗

Automated Immunoprecipitation Workflow for Comprehensive Acetylome Analysis

Immunoprecipitation is one of the most effective methods for enrichment of lysine-acetylated peptides for comprehensive acetylome analysis using mass spectrometry. Manual acetyl peptide enrichment method using non-conjugated antibodies and agarose beads has been developed and applied in various studies. However, it is time consuming, and can introduce contaminants and variability that leads to potential sample loss and decreased sensitivity and robustness of the analysis. Here we describe a fast, automated enrichment protocol that enables reproducible and comprehensive acetylome analysis using a magnetic bead-based immunoprecipitation reagent.

Lysine acetylation, Acetylome, Acetyl peptide enri↗

HiFiAdapterFilt, a memory efficient read processing pipeline, prevents occurrence of adapter sequence in PacBio HiFi reads and their negative impacts on genome assembly

Abstract Background Pacific Biosciences HiFi read technology is currently the industry standard for high accuracy long-read sequencing that has been widely adopted by large sequencing and assembly initiatives for generation of de novo assemblies in non-model organisms. Though adapter contamination filtering is routine in traditional short-read analysis pipelines, it has not been widely adopted for HiFi workflows. Results Analysis of 55 publicly available HiFi datasets revealed that a read-sanitation step to remove sequence artifacts derived from PacBio library preparation from read pools is necessary as adapter sequences can be erroneously integrated into assemblies. Conclusions Here we describe the nature of adapter contaminated reads, their consequences in assembly, and present HiFiAdapterFilt, a simple and memory efficient solution for removing adapter contaminated reads prior to assembly.

59 BASIC BIOLOGICAL SCIENCES↗

Near real-time streaming analysis of big fusion data

Experiments on fusion plasmas produce high-dimensional data time series with ever-increasing magnitude and velocity, but turn-around times for analysis of this data have not kept up. For example, many data analysis tasks are often performed in a manual, ad-hoc manner some time after an experiment. In this article, we introduce the Delta framework that facilitates near real-time streaming analysis of big and fast fusion data. By streaming measurement data from fusion experiments to a high-performance compute center, Delta allows computationally expensive data analysis tasks to be performed in between plasma pulses. This article describes the modular and expandable software architecture of Delta and presents performance benchmarks of individual components as well as of an example workflow. Focusing on a streaming analysis workflow where electron cyclotron emission imaging (ECEi) data is measured at KSTAR on the National Energy Research Scientific Computing Center's (NERSC's) supercomputer we routinely observe data transfer rates of about 4 Gigabit per second. In NERSC, a demanding turbulence analysis workflow effectively utilizes multiple nodes and graphical processing units and executes them in under 5 min. We further discuss how Delta uses modern database systems and container orchestration services to provide web-based real-time data visualization. For the case of ECEi data we demonstrate how data visualizations can be augmented with outputs from machine learning models. Here, by providing session leaders and physics operators, results of higher-order data analysis using live visualizations may make more informed decisions on how to configure the machine for the next shot.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Apparatus and method for safety analysis evaluation with data-driven workflow

An apparatus and method for system safety analysis evaluation is provided, the apparatus including processing circuitry configured for generating a calculation matrix for a system, generating a plurality of models based on the calculation matrix, performing a benchmarking or convolution analysis of the plurality of models, identifying a design envelope based on the benchmarking or convolution analysis, deriving uncertainty models from the benchmarking or convolution analysis, deriving an assessment judgment based on the uncertainty models and acceptance criteria, defining one or more limiting scenarios based on the design envelope, and determining a safety margin in at least one figure-of-merit for the system based on the design envelope and the acceptance criteria.

Martin, Robert P.↗

Apparatus and method for safety analysis evaluation with data-driven workflow

An apparatus and method for system safety analysis evaluation is provided, the apparatus including processing circuitry configured for generating a calculation matrix for a system, generating a plurality of models based on the calculation matrix, performing a benchmarking or convolution analysis of the plurality of models, identifying a design envelope based on the benchmarking or convolution analysis, deriving uncertainty models from the benchmarking or convolution analysis, deriving an assessment judgment based on the uncertainty models and acceptance criteria, defining one or more limiting scenarios based on the design envelope, and determining a safety margin in at least one figure-of-merit for the system based on the design envelope and the acceptance criteria.

Martin, Robert P.↗