Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Echo: Data and Analysis Management [Slides]

Echo software capabilities are described. Echo provides a comprehensive suite of tools for data wrangling, management, and analysis. It is an object-oriented approach to data analysis designed for robust and scalable execution of complex analysis workflows written in MATLAB.

97 MATHEMATICS AND COMPUTING↗

Ion Mobility Spectrometry-Mass Spectrometry for High-Throughput Analysis

Ion mobility spectrometry is a widely used analytical technique providing gas phase separation of molecules. It has received increasing attention in the recent years with the advancement in technology development and the availability of commercial instruments. In this chapter, we introduced the ion mobility fundamental theory and provided examples of IMS applications, especially for isomer separation, collision cross section database generation, high throughput analysis workflows, software tools for IMS data analysis, and ongoing high resolution SLIM IMS development. While IMS is not yet routinely utilized in drug discovery and pharmaceutical industry, there has been increased interest in high throughput library screening and antibody characterization. With all the ongoing development in IMS technology and informatics, we foresee more and more exciting applications of high throughput IMS analysis in different fields including omics studies, drug discovery and clinical applications in the near future.

Ross, Dylan H.↗

Multibody for Everybody (M4E) - A Linearization Approach to Enable Frequency Domain Analysis, Time Integration and Control Co-Design

1.1 Background/Objectives: Marine energy represents a promising yet underexploited source of power. To increase the harvested power, significant efforts have been made to improve wave energy converter (WEC) modeling capabilities and optimize power take-off (PTO) performance; however, these efforts have often treated WEC dynamics, PTO design, and controller development sequentially. In contrast, control co-design (CCD) is emerging as a promising strategy to address these issues directly, creating a growing need for fast analysis tools suitable for repeated simulation and parametric studies [1]. To support this need, this work presents the Multibody for Everybody (M4E) [2] linearization module, which employs a symbolic toolbox to provide deeper insight of WEC design parameters. The objective is to demonstrate that a minimal-coordinate linearization of articulated WEC dynamics can provide accurate wave response predictions and substantial computational savings relative to nonlinear time-domain simulation, while preserving compatibility with broader wave-energy analysis workflows, enabling CCD. 1.2 Approach/Activities: The proposed approach linearizes the equations of motion, generated by M4E, in minimal coordinates about a selected operating point and combines the resulting system with frequencydomain hydrodynamic terms to incorporate the reduced mass, damping, stiffness, and forcing operators. The linearized model is used for both impedance-based response amplitude operator (RAO) prediction and rapid regular-wave time integration. The methodology is demonstrated on a single-flap device and a FOSWEC configuration, with linearized M4E responses compared against the corresponding nonlinear M4E simulations and WEC-Sim results. Regular-wave time histories, RAO trends, and runtime differences are assessed. The framework is also compatible with broader wave-energy workflows, including coupling to WecOptTool, although that capability is not the focus of this work [3]. 1.3 Results/Lessons: The linearized M4E model reproduces key regularwave response characteristics such as integration and Response Amplitude over multiple frequencies. This module matches nonlinear M4E and WEC-Sim results while substantially reducing integration cost. Thus, the proposed framework can serve as a rapid analysis layer for articulated WEC design, parameter studies, and controls-oriented workflows. The analysis is most appropriate in the near-equilibrium regime, about the linearization point.

16 TIDAL AND WAVE POWER↗

Performance Analysis and Optimization for Scientific Data Workloads

Scientific data generated at experimental and observational facilities are increasingly being processed on large-scale compute systems. Most of the experimental data analysis workflows are not designed or implemented to run on large scale environments and take full advantage of HPC compute and storage resources. These applications are unlike the traditional tightly-coupled scientific applications and hence face significant performance and scalability challenges as the volume of data increases exponentially. In this paper, we conduct a performance and scalability analysis for experimental analysis applications and workflows operating on data from light sources. Our analysis detects and quantifies I/O performance, scalability and runtime bottlenecks for three data analysis applications that run on NERSC resources. Based on our analysis we propose and implement a set of optimizations that lead to reducing the amount of time spent on I/O operations by almost 90%.

97 MATHEMATICS AND COMPUTING↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

An intelligent Data Delivery Service for and beyond the ATLAS experiment

The intelligent Data Delivery Service (iDDS) has been developed to cope with the huge increase of computing and storage resource usage in the coming LHC data taking. It has been designed to intelligently orchestrate workflows and data management systems, decoupling data pre-processing, delivery, and primary processing in large scale workflows. It is an experiment-agnostic service that has been deployed to serve data carousel (orchestrating efficient processing of tape-resident data), machine learning hyperparameter optimization, active learning, and other complex multi-stage workflows defined via DAG (Directed Acyclic Graph), CWL (Common Workflow Language) and other descriptions, including a growing number of analysis workflows. We will at first introduce some deployed use cases in a summary. Then we will focus on new improvements and use cases under developments in ATLAS, Rubin Observatory and sPHENIX, together with future efforts.

97 MATHEMATICS AND COMPUTING↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Q3 Report for FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment

This report describes the work and activities carried out towards the completion of each of the following milestones in FY25 Q3: 1. Demonstrate workflow for generating self-consistent CESOL plasma profiles + first wall and divertor loading prediction and generate the CAT plasma and neutron loading needed for further engineering analysis. • Benchmark between two first wall heat flux mapping methods, identify importance of various heat flux sources and physics impact of using fully coupled CESOL vs post-analysis evaluation. 2. Generate medium fidelity parametrized CAD. • Generate parametrized CAD components for the CAT example case via either user-defined modules called within the geometry generation or by defeatured/parametrized CAD, including DCLL blanket matched to divertor boundary and magnets. Define materials, labels, and boundary conditions for passing the mesh to CFD tools. 3. Demonstrate multiphysics magnet analysis. • Demonstrate magnet analysis workflow called from the FREDA workflow, and 4. Demonstrate nuclear analysis. • Add model to OpenFOAM and/or other codes possibly including Diablo to account for tritium diffusion in solids.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

Nanometer Scale Imaging to Develop Quantitative Descriptors of Bipolar Membrane Junction Structure

Swings in pH can be achieved by electrically polarizing a bipolar membrane (BPM) to drive water dissociation at the BPM junction for electrochemical conversion and separation processes. BPM junction design is critical to tailor performance for specific applications; however, characterization techniques capable of resolving the nanometer scale physical structure of the junction are limited. We present sample preparation, imaging, and analysis workflows that are adaptable to a variety of BPM junction architectures. Atomic force microscopy produces BPM junction images with nanometer scale lateral resolution for samples with and without a graphene oxide water dissociation catalyst in the junction. Subsequent image segmentation and analysis quantify line edge roughness and catalyst layer thickness as descriptors of junction structure. Comparison of pre- and post-electrodialysis junctions suggests electric field-induced alignment of catalyst particles during electrodialysis. This characterization workflow can inform manufacturing protocols, computational modeling, and failure mode analysis for next-generation BPMs.

97 MATHEMATICS AND COMPUTING↗

Osprey Framework v0.2.2

The Alpha Berkeley Framework is a software architecture for building agentic AI systems that coordinate multi-step workflows in scientific and industrial environments. It is based on a plan-first orchestration model, where natural language requests are translated into execution plans with explicit dependencies and optional human approval. The framework includes capability classification, which selects relevant tools on a per-task basis to keep orchestration efficient as the number of available tools grows. It incorporates task extraction methods that compress conversational context and integrate external resources such as databases, APIs, and knowledge bases into structured, machine-readable tasks. Execution is supported by modular services with checkpointing, artifact management, and error handling, allowing workflows to be paused, inspected, and resumed. The system is designed for deployment in production environments, supporting both local and containerized execution as well as integration with HPC clusters. Interfaces include command-line tools, browser-based workflows, and containerized services. The framework has been demonstrated in tutorial examples and deployed at the Advanced Light Source, where it coordinates accelerator control and analysis workflows.

Hellert, Thorsten [Lawrence Berkeley National Labo↗

A galactic approach to neutron scattering science

Neutron scattering science is leading to significant advances in our understanding of materials and will be key to solving many of the challenges that society is facing today. Improvements in scientific instruments are actually making it more difficult to analyze and interpret the results of experiments due to the vast increases in the volume and complexity of data being produced and the associated computational requirements for processing that data. New approaches to enable scientists to leverage computational resources are required, and Oak Ridge National Laboratory (ORNL) has been at the forefront of developing these technologies. We recently completed the design and initial implementation of a neutrons data interpretation platform that allows seamless access to the computational resources provided by ORNL. For the first time, we have demonstrated that this platform can be used for advanced data analysis of correlated quantum materials by utilizing the world's most powerful computer system, Frontier. In particular, we have shown the end-to-end execution of the DCA++ code to determine the dynamic magnetic spin susceptibility χ(q, ω) for a single-band Hubbard model with Coulomb repulsion U/t = 8 in units of the nearest-neighbor hopping amplitude t and an electron density of n = 0.65. The following work describes the architecture, design, and implementation of the platform and how we constructed a correlated quantum materials analysis workflow to demonstrate the viability of this system to produce scientific results.

97 MATHEMATICS AND COMPUTING↗

Materials characterization: Can artificial intelligence be used to address reproducibility challenges?

Material characterization techniques are widely used to characterize the physical and chemical properties of materials at the nanoscale and, thus, play central roles in material scientific discoveries. However, the large and complex datasets generated by these techniques often require significant human effort to interpret and extract meaningful physicochemical insights. Artificial intelligence (AI) techniques such as machine learning (ML) have the potential to improve the efficiency and accuracy of surface analysis by automating data analysis and interpretation. In this perspective paper, we review the current role of AI in surface analysis and discuss its future potential to accelerate discoveries in surface science, materials science, and interface science. We highlight several applications where AI has already been used to analyze surface analysis data, including the identification of crystal structures from XRD data, analysis of XPS spectra for surface composition, and the interpretation of TEM and SEM images for particle morphology and size. We also discuss the challenges and opportunities associated with the integration of AI into surface analysis workflows. These include the need for large and diverse datasets for training ML models, the importance of feature selection and representation, and the potential for ML to enable new insights and discoveries by identifying patterns and relationships in complex datasets. Most importantly, AI analyzed data must not just find the best mathematical description of the data, but it must find the most physical and chemically meaningful results. In addition, the need for reproducibility in scientific research has become increasingly important in recent years. The advancement of AI, including both conventional and the increasing popular deep learning, is showing promise in addressing those challenges by enabling the execution and verification of scientific progress. By training models on large experimental datasets and providing automated analysis and data interpretation, AI can help to ensure that scientific results are reproducible and reliable. Although integration of knowledge and AI models must be considered for the transparency and interpretability of models, the incorporation of AI into the data collection and processing workflow will significantly enhance the efficiency and accuracy of various surface analysis techniques and deepen our understanding at an accelerated pace.

Materials Science↗

The Artificial Scientist: in-Transit Machine Learning of Plasma Simulations

Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.

Kelling, Jeffrey [Helmholtz-Zentrum Dresden Rossen↗

Seeking regularity from irregularity: unveiling the synthesis–nanomorphology relationships of heterogeneous nanomaterials using unsupervised machine learning

Nanoscale morphology of functional materials determines their chemical and physical properties. However, despite increasing use of transmission electron microscopy (TEM) to directly image nanomorphology, it remains challenging to quantify the information embedded in TEM data sets, and to use nanomorphology to link synthesis and processing conditions to properties. We develop an automated, descriptor-free analysis workflow for TEM data that utilizes convolutional neural networks and unsupervised learning to quantify and classify nanomorphology, and thereby reveal synthesis–nanomorphology relationships in three different systems. While TEM records nanomorphology readily in two-dimensional (2D) images or three-dimensional (3D) tomograms, we advance the analysis of these images by identifying and applying a universal shape fingerprint function to characterize nanomorphology. After dimensionality reduction through principal component analysis, this function then serves as the input for morphology grouping through unsupervised learning. We demonstrate the wide applicability of our workflow to both 2D and 3D TEM data sets, and to both inorganic and organic nanomaterials, including tetrahedral gold nanoparticles mixed with irregularly shaped impurities, hybrid polymer-patched gold nanoprisms, and polyamide membranes with irregular and heterogeneous 3D crumple structures. In each of these systems, unsupervised nanomorphology grouping identifies both the diversity and the similarity of the nanomaterial across different synthesis conditions, revealing how synthetic parameters guide nanomorphology development. Our work opens possibilities for enhancing synthesis of nanomaterials through artificial intelligence and for understanding and controlling complex nanomorphology, both for 2D systems and in the far less explored case of 3D structures, such as those with embedded voids or hidden interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TR-XPS Realtime Analysis Tool (ArroyoXPS) v0.1

The ALS has developed a Time-Resolved X-ray Photoelectron Spectroscopy (TR-XPS) technique, which involves applying a specific pattern of voltage curves to a sample while measuring XPS peaks. This pattern is repeated over multiple cycles, and changes in the material's response provide valuable scientific insights. Traditionally, file-based analysis workflows have been used: scans are run for a predetermined time, and after one or more scans are complete, calculations are made. ArroyoXPS changes this by offering in-experiment scan and analysis, allowing researchers to gain insights before a scan is finished. This enables them to adjust experimental parameters quickly, potentially saving valuable beamtime. ArroyoXPS includes tools for integrating with beamline control systems, performing analysis, and visualizing scan data in a web browser.

McReynolds, Dylan [Lawrence Berkeley National Labo↗