Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

A Systematic Interpretation of Subsurface Proppant Concentration from Drilling Mud Returns: Case Study from Hydraulic Fracturing Test Site (HFTS-2) in Delaware Basin

The aim of this study is generation and validation of a proppant log using analysis of drilling mud returns for child wells. Proppant log provides qualitative as well as quantitative insights into spatial distribution of proppant sand particles from prior stimulation of parent wells. While the basic methodology was developed and formalized during analysis of material collected from through fracture cores at Hydraulic Fracturing Test Site in Midland Basin (HFTS – 1), the test wells at HFTS – 2 in the neighboring Delaware Basin allowed the opportunity to validate the workflow on actual mud return samples from subsurface. As a child well is being drilled, periodic mud return samples are collected at the rig site and preserved for analysis. The workflow involves systematic cleaning of the samples including various steps such as washing, drying and segregation of samples into relevant size fractions of interest (< Mesh 20) based on specifications of pumped sand during stimulation of the parent well. Clean samples are imaged using high resolution transparency scanning. Scan images are then systematically analyzed for particles of interest using computer vision techniques. Sample counts are further validated using elemental analysis of smaller sub-samples at various depths of interest. This step is necessary to isolate proppant versus other naturally occurring minerals such as sulphates and carbonates which show similar optical properties. We successfully correlated proppant distribution against the existing parent well and validated propped versus relatively un-propped zones for a child well at the test site. The advantage of testing the proppant log concept at the HFTS – 2 site is the plethora of additional diagnostic data that is available to validate our primary observations. We can correlate spatial proppant distribution against variability in stimulation response based on independent observations such as image logs, microseismic attributes as well as DAS response, all of which tend to corroborate one another. One of our significant successes was being able to describe varying degrees of impact of the parent well along the lateral length of a stimulated child well. Our workflow represents a systematic and one-of-a-kind interpretation of spatial proppant distribution while drilling child wells. This provides unique opportunities to better understand the current state of the Downloaded from http://onepetro.org/URTECONF/proceedings-pdf/21URTC/2-21URTC/D021S031R003/2477415/urtec-2021-5189-ms.pdf/1 by Carol Worster on 28 February 2022 URTeC 5189 2 reservoir being targeted including zones which are likely more drained relative to others and how the planned completion of the child well can be improved. Lastly, this log can be useful is validating optimal well spacing in relatively new fields under development.

58 GEOSCIENCES↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool designed to both prototype and evaluate these coupled workflows. In this paper, we use SimAI-Bench to benchmark the data transport performance of two common patterns on the Aurora supercomputer: a one-to-one workflow with co-located simulation and AI training instances, and a many-to-one workflow where a single AI model is trained from an ensemble of simulations. For the one-to-one pattern, our analysis shows that node-local and DragonHPC data staging strategies provide excellent performance compared Redis and Lustre file system. For the many-to-one pattern, we find that data transport becomes a dominant bottleneck as the ensemble size grows. Our evaluation reveals that file system is the optimal solution among the tested strategies for the many-to-one pattern.

Tummalapalli, Harikrishna [Argonne National Labora↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗

Darshan for HEP applications

Modern HEP workflows must manage increasingly large and complex data collections. HPC facilities may be employed to help meet these workflows’ growing data processing needs. However, a better understanding of the I/O patterns and underlying bottlenecks of these workflows is necessary to meet the performance expectations of HPC systems.Darshan is a lightweight I/O characterization tool that captures concise views of HPC application I/O behavior. It intercepts application I/O calls at runtime, records file access statistics for each process, and generates log files detailing application I/O access patterns.Typical HEP workflows include event generation, detector simulation, event reconstruction, and subsequent analysis stages. A study of the I/O behavior of the ATLAS simulation and filtering stage, and the CMS simulation workflow using Darshan is presented, including insights into the I/O operations and data access size.

Wang, Rui↗

Cognitive analysis of metabolomics data for systems biology

Cognitive computing is revolutionizing the way big data are processed and integrated, with artificial intelligence (AI) natural language processing (NLP) platforms helping researchers to efficiently search and digest the vast scientific literature. Most available platforms have been developed for biomedical researchers, but new NLP tools are emerging for biologists in other fields and an important example is metabolomics. NLP provides literature-based contextualization of metabolic features that decreases the time and expert-level subject knowledge required during the prioritization, identification and interpretation steps in the metabolomics data analysis pipeline. Here, we describe and demonstrate four workflows that combine metabolomics data with NLP-based literature searches of scientific databases to aid in the analysis of metabolomics data and their biological interpretation. Additionally, the four procedures can be used in isolation or consecutively, depending on the research questions. The first, used for initial metabolite annotation and prioritization, creates a list of metabolites that would be interesting for follow-up. The second workflow finds literature evidence of the activity of metabolites and metabolic pathways in governing the biological condition on a systems biology level. The third is used to identify candidate biomarkers, and the fourth looks for metabolic conditions or drug-repurposing targets that the two diseases have in common. The protocol can take 1–4 h or more to complete, depending on the processing time of the various software used.

59 BASIC BIOLOGICAL SCIENCES↗

Ensemble learning-iterative training machine learning for uncertainty quantification and automated experiment in atom-resolved microscopy

Deep learning has emerged as a technique of choice for rapid feature extraction across imaging disciplines, allowing rapid conversion of the data streams to spatial or spatiotemporal arrays of features of interest. However, applications of deep learning in experimental domains are often limited by the out-of-distribution drift between the experiments, where the network trained for one set of imaging conditions becomes sub-optimal for different ones. This limitation is particularly stringent in the quest to have an automated experiment setting, where retraining or transfer learning becomes impractical due to the need for human intervention and associated latencies. Here we explore the reproducibility of deep learning for feature extraction in atom-resolved electron microscopy and introduce workflows based on ensemble learning and iterative training to greatly improve feature detection. This approach allows incorporating uncertainty quantification into the deep learning analysis and also enables rapid automated experimental workflows where retraining of the network to compensate for out-of-distribution drift due to subtle change in imaging conditions is substituted for human operator or programmatic selection of networks from the ensemble. This methodology can be further applied to machine learning workflows in other imaging areas including optical and chemical imaging.

36 MATERIALS SCIENCE↗

Technical note: Optimizing the in situ cosmogenic 36 Cl extraction and measurement workflow for geologic applications

Abstract. In situ cosmogenic 36Cl analysis by accelerator mass spectrometry (AMS) is routinely employed to date Quaternary surfaces and assess rates of landscape evolution. However, standard laboratory preparation procedures for 36Cl dating require the addition of large amounts of isotopically enriched chlorine spike solution; these solutions are expensive and increasingly difficult to acquire from commercial sources. In addition, the typical workflow for 36Cl dating involves measuring both 35Cl/37Cl and 36Cl/Cl concurrently on the high-energy (post-accelerator) end of the AMS system, but 35Cl/37Cl determinations using this technique can be complicated by isotope fractionation and system memory during measurement. The traditional workflow also does not provide 36Cl extraction laboratories with the data needed to calculate native Cl concentrations in advance of 36Cl/Cl measurements. In light of these concerns, we present an improved workflow for extracting and measuring chlorine in geologic materials. Our initial step is to characterize 35Cl/37Cl on sample aliquots of up to ∼1 g prepared in Ag(Cl, Br) matrices, which greatly reduces the amount of isotopically enriched spike solution required to measure native Cl content in each sample. To avoid potential issues with isotope fractionation through the accelerator, 35Cl/37Cl is measured on the low-energy, pre-accelerator end of the AMS line. Then, for 36Cl/Cl measurements, we extract Cl as AgCl or Ag(Cl, Br) in analytical batches with a consistent total Cl load across all samples; this step is intended to minimize source memory effects during 36Cl/Cl measurements and allows the preparation of AMS standards that are customized to match known Cl contents in the samples. To assess the efficacy of this extraction and measurement workflow, we compare chlorine isotope ratio measurements on seven geologic samples prepared using standard procedures and the updated workflow. Measurements of 35Cl/37Cl and 36Cl/Cl are consistent between the two workflows, and 35Cl/37Cl values measured using our methods have considerably higher precision than those measured following standard protocols. The chemical preparation and measurement workflow presented here (1) reduces the amount of isotopically enriched chlorine spike used per rock sample by up to 95 %; (2) identifies rocks with high native Cl concentrations, which may be lower priority for 36Cl surface exposure dating, at an early stage of analysis; and (3) allows laboratory users to maintain control over the total chlorine content within and across analytical batches. These methods can be incorporated into existing laboratory and AMS protocols for 36Cl analyses and will increase the accessibility of 36Cl dating for geologic applications.

58 GEOSCIENCES↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

An Adaptive Geometry-Free Thermo-Mechanical Model for Directed Energy Deposition Process Modeling

This presentation describes a novel, geometry-free thermo-mechanical model with adaptive subdomain con- struction to accurately predict the thermal conditions, distortions, and residual stresses throughout the directed energy deposition (DED) process. A novel finite element workflow is designed to con- duct the numerical analysis, based on the multi-app and data transfer capabilities in the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE). Unlike with traditional methods, the part geometry in this model is not predefined. Instead, it is a combined effect of the processing parameters and material properties. At each time step, the model utilizes a subdomain construction paradigm to model the material deposition. A specialized mesh adaptivity scheme is incorporated to provide an accurate prediction while reducing the overall computational cost. The results generated by the proposed model show general agreement with the experimental measurements for the single track scan with varying processing parameters and demonstrate reasonable predictions for higher material buildups.

36 MATERIALS SCIENCE↗

Q1 Report for FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment

This report describes the work and activities carried out towards the completion of each of the following milestones in FY25 Q1: 1. Demonstrate workflow for generating self-consistent CESOL plasma profiles + first wall and divertor loading prediction and generate the CAT plasma and neutron loading needed for further engineering analysis: $/circ$ Initially run CESOL with SOLPS and map heat flux using simple HEAT method: ▪ Generate SOLPS grid for CAT reference case, ▪ Implement and apply ‘lower’ fidelity method of using HEAT-like analytic method to map SOLPS heat flux from charged particles to first wall, and ▪ Provide initial heat flux + neutron loading for further engineering analysis. 2. Demonstrate multiphysics magnet analysis: $/circ$ Demonstrate validation of the Elmer workflow by comparing induced stresses due to EM forces generated by TF coils using multiple codes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Coupled Target-Beam-Moderator Optimization for the Second Target Station

This report describes the results for a coupled target-beam-moderator optimization analysis for the Second Target Station (STS) at ORNL's Spallation Neutron Source. This study is a continuation of the optimization analysis for the moderators in the preliminary design of STS performed in 2022. In the 2022 analysis the dimensions of the moderators are parameterized, while the target and the proton beam profile are kept constant. In this analysis the target height and the proton beam profile are added as parameters. This allows to study the coupled effects of changing target, moderator and beam dimensions. Similar to the 2022 analysis, this work is performed with an automated optimization workflow that uses the optimization toolbox DAKOTA, parameterized geometries in CREO and SpaceClaim, the unstructured mesh generation in Attila4MC, and the particle transport code MCNP6.2©. This workflow enables an efficient optimization using high-fidelity geometries. The main conclusions of this analysis are the following: • Coupled beam-target-moderator optimization provides a few additional percent performance gain over stand-alone moderator optimization. • The moderator performance is not very sensitive to the target height (between ≈60 and ≈80 mm) as long as the beam profile is chosen adequately. • The moderator performance is sensitive to the choice of beam spatial standard deviations, even when the footprint is kept constant. • The optimal moderator radius is the same for a beam footprint of 30 cm 2 , 62.5 cm 2 , and 90 cm 2 . Also the slope of the super-gaussian beam profile does not significantly impact the optimal moderator radius. • The optimal parameters and sensitivities are very similar to the 2022 optimization analysis. These results only indicate a a difference in the optimal radius of the cylindrical moderator, however, this has been corrected in the final design moderator optimization. The main purpose of this report is to document the simulations, results and lessons learned. The most impactful results are summarized in. We also note that the target geometry used in this work is not the final design.

43 PARTICLE ACCELERATORS↗

Algorithms and file structures to enhance software workflows for ion mobility mass spectrometry (IM-MS)

Support customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based metabolite characterization. Evaluate and improve the integration of ion mobility to existing MS analysis methods of the Mass Profiler Professional workflow (Mass Profiler, ID Browser and Mass Profiler Professional).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

3D high-fidelity automated neutronics guided optimization of fusion blanket designs

The compact Fusion Pilot Plant (FPP) is defined in the recent National Academies of Sciences, Engineering, and Medicine report as the next step of fusion energy demonstration with a $50$ MWe peak net electricity production, $Q_e$ greater than $1$, and at least $3$ hours of continuous operation. This fusion pilot plant will be a test bed enabling materials, designs, and fuel management assessment, and it will represent an engineering challenge because of its high-fusion power and compact design targets. Previous reactor data is limited to experiments operating in different design space ranges. Therefore, design iterations and assessments should rely on high-fidelity first-principle theoretical and computational models. The high-fidelity integrated modeling of the plasma is a fundamental part of fusion energy research. However, the whole device modeling is often neglected, utilizing low-fidelity, system-level analysis. Recently, the need for high-fidelity multi-physics modeling was recognized, resulting in a selection of integrated tools. Further, autonomous design optimization requires a streamlined framework that perturbs the design point, reruns the analysis, and examines the outputs. However, high-fidelity analysis requires complex geometry specification that is difficult to perturb. This work presents the parametric CAD generation tool TRACER and a new neutronic workflow. TRACER allows the perturbation of the geometry representation, creating geometry files ready for further analysis. The streamlined neutronic workflow allows efficient and accurate calculations. The two new tools coupled together were used to perform a 3D high-fidelity multi-objective, multi-input optimization of an "ARC Class" compact tokamak design. The workflow was driven by an optimization driver for full automation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Image-Based Methods to Score Fungal Pathogen Symptom Progression and Severity in Excised Arabidopsis Leaves

Image-based symptom scoring of plant diseases is a powerful tool for associating disease resistance with plant genotypes. Advancements in technology have enabled new imaging and image processing strategies for statistical analysis of time-course experiments. There are several tools available for analyzing symptoms on leaves and fruits of crop plants, but only a few are available for the model plant Arabidopsis thaliana (Arabidopsis). Arabidopsis and the model fungus Botrytis cinerea (Botrytis) comprise a potent model pathosystem for the identification of signaling pathways conferring immunity against this broad host-range necrotrophic fungus. Here, we present two strategies to assess severity and symptom progression of Botrytis infection over time in Arabidopsis leaves. Thus, a pixel classification strategy using color hue values from red-green-blue (RGB) images and a random forest algorithm was used to establish necrotic, chlorotic, and healthy leaf areas. Secondly, using chlorophyll fluorescence (ChlFl) imaging, the maximum quantum yield of photosystem II (Fv/Fm) was determined to define diseased areas and their proportion per total leaf area. Both RGB and ChlFl imaging strategies were employed to track disease progression over time. This has provided a robust and sensitive method for detecting sensitive or resistant genetic backgrounds. A full methodological workflow, from plant culture to data analysis, is described.

59 BASIC BIOLOGICAL SCIENCES↗

Model Diagnostics for Equation Oriented Models: Roadblocks and the Path Forward

This poster was presented at the Foundations of Computer Aided Process Design (FOCAPD) 2024 conference with a paper describing efforts to develop a unified workflow and toolbox for diagnosing issues in equation-oriented models as part of the IDAES modeling framework. The poster presents the work currently underway within the IDAES project to develop and integrate cutting edge model analysis techniques into a common toolbox and workflow for model developers to use.

Lee, Andrew↗