Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Synergizing human expertise and AI efficiency with language model for microscopy operation and automated experiment design

With the advent of large language models (LLMs), in both the open source and proprietary domains, attention is turning to how to exploit such artificial intelligence (AI) systems in assisting complex scientific tasks, such as material synthesis, characterization, analysis and discovery. Here, we explore the utility of LLMs, particularly ChatGPT4, in combination with application program interfaces (APIs) in tasks of experimental design, programming workflows, and data analysis in scanning probe microscopy, using both in-house developed APIs and APIs given by a commercial vendor for instrument control. We find that the LLM can be especially useful in converting ideations of experimental workflows to executable code on microscope APIs. Beyond code generation, we find that the GPT4 is capable of analyzing microscopy images in a generic sense. At the same time, we find that GPT4 suffers from an inability to extend beyond basic analyses for more in-depth technical experimental design. We argue that an LLM specifically fine-tuned for individual scientific domains can potentially be a better language interface for converting scientific ideations from human experts to executable workflows. Such a synergy between human expertise and LLM efficiency in experimentation can open new doors for accelerating scientific research, enabling effective experimental protocols sharing in the scientific community.

97 MATHEMATICS AND COMPUTING↗

DaYu: Optimizing Distributed Scientific Workflows by Decoding Dataflow Semantics and Dynamics

The combination of ever-growing scientific datasets and distributed workflow complexity creates I/O performance bottlenecks due to data volume, velocity, and variety. Although the increasing use of descriptive data formats (e.g., HDF5, netCDF) helps organize these datasets, it also creates obscure bottlenecks due to the need to translate high level operations into file addresses and then into low-level I/O operations. To address this challenge, we introduce DaYu, a method and toolset for analyzing (a) semantic relationships between logical datasets and file addresses, (b) how dataset operations translate into I/O, and (c) the combination across entire workflows. DaYu's analysis and visualization enables identification of critical bottlenecks and reasoning about remediation. We describe our methodology and propose optimization guidelines. Evaluation on scientific workflows demonstrates up to 3.7x performance improvements in I/O time for obscure bottlenecks. The time and storage overhead for DaYu's time-ordered data is typically under 0.2% of runtime and 0.25% of data volume, respectively.

Tang, Meng↗

An automatically curated first-principles database of ferroelectrics

Ferroelectric materials have technological applications in information storage and electronic devices. The ferroelectric polar phase can be controlled with external fields, chemical substitution and size-effects in bulk and ultrathin film form, providing a platform for future technologies and for exploratory research. In this work, we integrate spin-polarized density functional theory (DFT) calculations, crystal structure databases, symmetry tools, workflow software, and a custom analysis toolkit to build a library of known, previously-proposed, and newly-proposed ferroelectric materials. With our automated workflow, we screen over 67,000 candidate materials from the Materials Project database to generate a dataset of 255 ferroelectric candidates, and propose 126 new ferroelectric materials. We benchmark our results against experimental data and previous first-principles results. The data provided includes atomic structures, output files, and DFT values of band gaps, energies, and the spontaneous polarization for each ferroelectric candidate. Finally, we contribute our workflow and analysis code to the open-source python packages atomate and pymatgen so others can conduct analogous symmetry driven searches for ferroelectrics and related phenomena.

36 MATERIALS SCIENCE↗

Simplifying Geospatial Workflows with GeoGridFusion

Growing demands to understand PV deployment and reliability across an expanding range of climates, along with increasing computational power, are driving the need for simplified computing tools that support large-scale analysis and geospatial workflows. While existing open-source libraries and PV system modeling tools offer extensive collections of empirical and physical models, they often lack the ability to scale to large geospatial datasets. The tool presented here, GeoGridFusion, enables PV modelers to store and utilize gridded geospatial data from sources such as NSRDB, PVGIS, and others. GeoGridFusion harmonizes diverse datasets, taxonomies, and nomenclatures, and includes utilities that support intuitive geospatial area selection.

14 SOLAR ENERGY↗

Advancing subsurface analysis: Integrating computer vision and deep learning for the near real-time interpretation of borehole image logs in the Illinois Basin-Decatur Project

The accurate quantification and mapping of subsurface natural fracture systems using borehole imaging logs are critical for the success of CO 2 sequestration in geologic formations, optimization of engineered geothermal systems, and hydrocarbon production enhancement. However, traditional interpretation processes suffer from time-consuming procedures and human bias. To address these challenges and expedite fracture analysis, we investigated the application of integrated computer vision and DL workflows to automate image log analysis. Specifically, the design of our workflow was crafted to swiftly detect fractures and baffles by using actual electrical resistivity of borehole wall from microresistivity imaging device alongside their binary representation. This novel approach significantly reduces computational time while providing invaluable insights. By incorporating conventional logging and microseismic data, we present a regional subsurface natural fracture mapping technique. Through the minimization of human bias in image log analysis, our automated workflow achieves reduced fracture interpretation time and costs while ensuring robust and reproducible results. We demonstrated the efficacy of our approach by applying the workflow to the Illinois Basin-Decatur Project site. The automated workflow successfully identified major fractured zones, multiple baffles, and an interbedded layer with a high resolution of 0.01 ft or 0.12 in. (0.3 cm) and can be upscaled to any desired resolution. Validation through microseismic and image log interpretations allows for accurate and near-real-time mapping of fractures and baffles, significantly enhancing CO 2 pressure forecasting and postinjection site care. Our approach stands out due to its robustness, consistency, and reduced computational cost compared with alternative feature extraction technologies. It presents exciting possibilities for advancing CO 2 sequestration and engineered geothermal efforts by offering comprehensive and efficient fracture mapping solutions. This technology can contribute significantly to the optimization of CO 2 sequestration projects, facilitating sustainable environmental practices, and combating climate change.

Geochemistry & Geophysics↗

FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment (Final Report)

This report documents the FY25 Theory and Simulation Performance Target (TSPT) of developing an integrated modeling framework for fusion reactor design and assessment (FREDA). Over Q1-Q4, new capabilities were developed across both plasma and engineering domains and demonstrated on an example representation of a Compact Advanced Tokamak with a Dual Cooled Lead Lithium blanket. This represents a first-of-a-kind demonstration of coupled core-to-wall-to-engineering for a reactor. Self-consistent CESOL workflows were applied to provide core, pedestal, and SOL prediction; new modules were developed for energetic particle stability (FAR3D) and transport (TGLF-EP) analysis; and boundary plasma modeling (SOLPS-ITER, BOUT++/Hermes-3) was expanded to evaluate wall and divertor heat fluxes and interface with engineering thermal analysis. A parameterized CAD tool, TRACER, was expanded to generate medium-fidelity divertor, blanket, and coil geometries; OpenFOAM and Diablo workflows were applied for first-wall and divertor thermal analyses with helium cooling; and reduced-order models were created for high-mass-flux divertor cooling. Magnet multiphysics capabilities were verified between Elmer, Diablo, and a new MFEM-based solver, and workflows enable stress, thermal, and neutron-fluence analysis of TF coils with neutronics-driven heating. Nuclear and blanket analysis workflows were demonstrated, including tritium breeding, transport, and CFD-informed thermo-mechanical assessment. Preliminary multi-fidelity uncertainty quantification workflows were applied to boundary modeling codes and shown to achieve variance reductions with fewer high-fidelity boundary simulations. Key findings highlight the challenges of resolving the ITEP gap to find suitable balance between wall and divertor loads, neutron heating, and practical limits of PFC cooling. Next step priorities are to develop automated workflows to check boundary code convergence and detachment, implement tighter physics-engineering CAD provenance tracking, and inclusion of plasma-material interface models for SLAG and tungsten cracking behavior. Collectively, these developments establish sophisticated capabilities for predictive, multi-fidelity, whole-device modeling that integrates plasma physics, materials, magnets, and nuclear engineering to guide pathways to viable Fusion Pilot Plant design points.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Toward Unified Autonomous Scattering Experiments: A Cross-Facility Case Study at ALS and PETRA III

Autonomous experiments rely on the integration of control, data acquisition, analysis, and decision-making frameworks. While such systems have been demonstrated at individual facilities, adapting them to additional instruments remains challenging due to differences in local infrastructure. We present a modular workflow that connects existing open-source tools for data access (Tiled), workflow orchestration (Prefect), analysis and visualization (pyFAI, Plotly Dash), and Gaussian-process-based adaptive sampling (gpCAM) into a unified framework for autonomous scattering experiments. The same configuration operates across two synchrotron beamlines (ALS 7.3.3 and PETRA III P03) with only minimal facility-specific adjustments, as shown in proof-of-concept demonstrations. This validates that a consistent design emphasizing modularity and shared interfaces can ease deployment across diverse experimental environments. The resulting framework provides a flexible foundation for extending autonomous control and analysis capabilities beyond a single beamline or instrument.

47 OTHER INSTRUMENTATION↗

An open source analysis framework for large-scale building energy modeling

Full integration of building energy modelling into the design and retrofit process has long been a goal of building scientists and practitioners. However, significant barriers still exist. Among them are the lack of available: (1) configurable technology stacks for performing both small- and large-scale analyses, (2) different classes of algorithms compatible with common design workflows, and (3) analysis tools for effectively visualizing large-scale simulation results. This article discusses the OpenStudio® Analysis Framework: a scalable analysis framework for building energy modelling that was developed to overcome the three barriers listed above. The framework is open-source and scalable to facilitate wider adoption and has a clearly defined application programming interface upon which other applications can be built. It runs on high-performance computing systems, within cloud infrastructure, and on laptops, and uses a common workflow to enable different classes of algorithms. Lessons learned from previous development efforts are also discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Benchmarking image processing techniques for porosity measurement in polymer additive manufacturing: Review and experimental analysis

An image processing workflow is proposed for porosity measurement in polymer additive manufacturing. Various techniques, including global and local thresholding, region growing, and K-means clustering, were applied to microscopic images of carbon fiber reinforced acrylonitrile butadiene styrene (CF-ABS) and benchmarked for their ability to accurately measure porosity. Global methods included Otsu, minimum error, iterative, and entropy-based thresholding, while local methods included Niblack, Bernsen, Sauvola, and Bradley-Roth algorithms. Artificial uneven illumination was introduced to test local adaptive thresholds. Results showed significant differences in porosity values across methods. Otsu, region growing, and K-means clustering excelled under uniform illumination, while Sauvola and Bradley-Roth performed better with uneven illumination. Comparison with X-ray computed tomography (XCT) revealed slightly lower porosity values (2.55 %) than optimized methods (2.73–2.79 %) due to XCT's lower resolution excluding smaller pores. While XCT offers finer pore detection, it limits sample volume and underestimates porosity due to spatial variation. Validation using artificial grayscale images with 5 % porosity confirmed that Otsu, Bradley-Roth, region growing, and Sauvola algorithms produced accurate results. Although tested on a single material system, these methods can be adapted to others with optimization. In conclusion, given XCT's high computational and time costs, this study highlights suitable image processing techniques as cost-effective alternatives for porosity analysis in polymer composites.

Additive manufacturing↗

Discovering invariant spatial features in electron energy loss spectroscopy images on the mesoscopic and atomic levels

Over the last two decades, Electron Energy Loss Spectroscopy (EELS) imaging with a scanning transmission electron microscope has emerged as a technique of choice for visualizing complex chemical, electronic, plasmonic, and phononic phenomena in complex materials and structures. The availability of the EELS data necessitates the development of methods to analyze multidimensional data sets with complex spatial and energy structures. Traditionally, the analysis of these data sets has been based on analysis of individual spectra, one at a time, whereas the spatial structure and correlations between individual spatial pixels containing the relevant information of the physics of underpinning processes have generally been ignored and analyzed only via the visualization as 2D maps. Here, we develop a machine learning-based approach and workflows for the analysis of spatial structures in 3D EELS data sets using a combination of dimensionality reduction and multichannel rotationally invariant variational autoencoders. This approach is illustrated for the analysis of both the plasmonic phenomena in a system of nanowires and in the core excitations in functional oxides using low loss and core-loss EELS, respectively. The code developed in this manuscript is open sourced and freely available and provided as a Jupyter notebook for the interested reader.

36 MATERIALS SCIENCE↗

Quantifying the Dynamics of Protein Self-Organization Using Deep Learning Analysis of Atomic Force Microscopy Data

The dynamics of protein self-assembly on the inorganic surface and the resultant geometric patterns are visualized using high-speed atomic force microscopy. The time dynamics of the classical macroscopic descriptors such as 2D fast Fourier transforms, correlation, and pair distribution functions are explored using the unsupervised linear unmixing, demonstrating the presence of static ordered and dynamic disordered phases and establishing their time dynamics. Here, the deep learning (DL)-based workflow is developed to analyze detailed particle dynamics and explore the evolution of local geometries. Finally, we use a combination of DL feature extraction and mixture modeling to define particle neighborhoods free of physics constraints, allowing for a separation of possible classes of particle behavior and identification of the associated transitions. Overall, this work establishes the workflow for the analysis of the self-organization processes in complex systems from observational data and provides insight into the fundamental mechanisms.

36 MATERIALS SCIENCE↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Near-Single-Cell Proteomics Profiling of the Proximal Tubular and Glomerulus of the Normal Human Kidney

Molecular assessments at the single-cell level can accelerate biological research by providing detailed assessments of cellular organization and tissue heterogeneity in both disease and health. The human kidney has complex multi-cellular states with varying functionality, much of which can now be completely harnessed with recent technological advances in tissue proteomics at a near single-cell level. We discuss the foundational steps in the first application of this mass spectrometry (MS) based proteomics method for analysis of sub-sections of the normal human kidney, as part of the Kidney Precision Medicine Project (KPMP). Using approximately 30-40 laser captured micro-dissected kidney cells, we identified more than 2500 human proteins, with specificity to the proximal tubular (PT; n=25 proteins) and glomerular (Glom; n=67 proteins) regions of the kidney and their unique metabolic functions. This pilot study provides the roadmap for application of our near-single-cell proteomics workflow for the analysis of other renal micro-compartments, on a larger scale, to unravel perturbations of renal sub-cellular function in the normal kidney as well as different etiologies of acute and chronic kidney disease.

60 APPLIED LIFE SCIENCES↗

Correlating and Simulating Socio-Demographically Driven Residential End-Use Activity Schedules

Incorporating socio-demographic and behavioral considerations into decision-support tools is crucial for identifying gaps and addressing consumer needs to ensure reliable and affordable energy solutions. In energy simulation models, the correlation between socio-demographics and time-use behavior is not well-captured. Thus, we developed a large-scale simulation workflow to generate schedules for 10 residential activities across 24 population segments defined by age, income, and employment status. Using pre-pandemic 2015-2019 American Time Use Survey (ATUS) data, we used ANOVA to confirm the correlation between demographic factors and time use. We explored three k-modes clustering methods-backward, forward, and a new hybrid approach-to delineate the occupancy patterns based on demographics. Using the probability of cluster membership for each population segment and a time inhomogeneous Markov chain to generate activity transition probabilities for each cluster, we simulated 50,000 schedules per segment and validated them against the ATUS data. The hybrid method produced the most socio-demographically differentiated clusters while demonstrating comparable performance to other approaches, with an overall root mean square error of 0.12 for both weekday and weekend schedules. Thus, the hybrid method, where each cluster is dominated by certain demographic segments and occupancy patterns, offers more modeling versatility in terms of scenario analysis. The new workflow improves the socio demographic differentiation of energy consumption by considering differences in time use. This approach enables future research on demographically segmented time of use (TOU) energy consumption, including impacts of TOU utility bills and rate analysis, long-run marginal emissions, and energy retrofits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Boosting RDataFrame performance with transparent bulk event processing

RDataFrame is ROOT’s high-level interface for Python and C++ data analysis. Since it first became available, RDataFrame adoption has grown steadily and it is now poised to be a major component of analysis software pipelines for LHC Run 3 and beyond. Thanks to its design inspired by declarative programming principles, RDataFrame enables the development of highperformance, highly parallel analyses without requiring expert knowledge of multi-threading and I/O: user logic is expressed in terms of self-contained, small computation kernels tied together by a high-level API. This design completely decouples analysis logic from its actual execution, and opens several interesting avenues for workflow optimization. In particular, in this work we explore the benefits of moving internal data processing from an event-by-event to a bulkby-bulk loop. This refactoring dramatically reduces the framework’s runtime overheads; in collaboration with the I/O layer it improves data access patterns; it exposes information that optimizing compilers might use to auto-vectorize the invocation of user-defined computations; finally, while existing user-facing interfaces remain unaffected, it becomes possible to additionally offer interfaces that explicitly expose bulks of events, useful e.g. for the injection of GPU kernels into the analysis workflow. In order to inform similar future R&D, design challenges will be presented, as well as an investigation of the relevant timememory trade-off backed by novel performance benchmarks.

Guiraud, Enrico↗

The ATLAS Workflow Management System Evolution in the LHC Run3 and towards the High-Luminosity LHC era

The ATLAS experiment has 18+ years of experience using workload management systems to deploy and develop workflows to process and to simulate data on the distributed computing infrastructure. Simulation, processing and analysis of LHC experiment data require the coordinated work of heterogeneous computing resources. In particular, the ATLAS experiment utilizes the resources of 250 computing centers worldwide, the power of supercomputing centres, and national, academic and commercial cloud computing resources. In this contribution, we present new techniques for cost-effectively improving efficiency introduced in workflow management system software. The evolution from a mesh framework to new types of computing facilities such as cloud and HPCs is described, as well as new types of production and analysis workflows.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Veridical data science

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, composed of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. As part of the PCS workflow, we develop PCS inference procedures, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others. Moreover, we demonstrate its favorable performance over existing methods in terms of receiver operating characteristic (ROC) curves in high-dimensional, sparse linear model simulations, including a wide range of misspecified models. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

97 MATHEMATICS AND COMPUTING↗

A New Workflow of X-ray CT Image Processing and Data Analysis of Structural Features in Rock Using Open-Source Software

X-ray computed tomography (CT) images of rock specimens often contain artifacts which must be corrected before scientific analyses are performed. Here, we present a new workflow of automated image processing to utilize poor-quality X-ray CT scan images. The workflow runs on the open-source image analysis software and efficiently separates desired features from low-contrast scanned images. The new workflow is a two-step technique using contrast enhancement and automated feature segmentation to generate noise-free binary images. The results of binary images using the proposed workflow and using a conventional thresholding technique are analyzed to show the quality of the proposed method. The paper also presents a workflow of estimating the structural geometries of features in two and three dimensions. The results of the structural feature analyses and computational time were compared between the open-source (ImageJ) and commercial image analysis software (Bruker Computed Tomography Analyzer). The commercial software was more computationally efficient, but the task-specific macros in open-source software enabled the user-desired automation in image processing and data extraction of desired structural features of comparable quality.

47 OTHER INSTRUMENTATION↗