Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Accessible Content Optimization for Research Needs (ACORN)

ACORN employs a set of automated processes for informing and/or enforcing defined content schemas to create standardized and highly structured data. Because of its standardized data source, ACORN easily applies computer automation to generate communication assets such as PDFs, Powerpoint presentations, and web pages. Built using the memory-safe Rust programming language, ACORN is portable and accessible for use on any Windows, Mac, or Linux machine.

Wohlgemuth, JasonHoward [Oak Ridge National Labora↗

Simulation process and data flow for a large system dynamics model

This paper documents the workflow and supporting technologies that a large system dynamics model, the biomass scenario model, employs to streamline the data preparation, simulation, quality control, and analysis process at the National Renewable Energy Laboratory. The workflow centers on automation of routine aspects of the flow of data between data stores, simulations, and visualizations. It enforces quality checks on data, reproducibility of computations, and traceability of results, while maintaining complete archives of modeling and analysis artifacts. The resulting frictionless simulation/analysis environment supports large-scale sensitivity analysis, interactive creation of ensembles of simulations, and rapid visualization-based exploration of simulation results.

09 BIOMASS FUELS↗

Integrated Risk-Informed Condition Based Maintenance Capability and Automated Platform: Technical Report 3

This project is a collaborative research effort between PKMJ Technical Services LLC, Idaho National Laboratory, and Public Service Enterprise Group (PSEG) Nuclear, LLC. The collaboration, led by PKMJ Technical Services LLC, is part of the industry Funding Opportunity Announcement (FOA) award under Advanced Nuclear Technology Development FOA #DE-FOA-0001817. The pilot demonstration focuses on the Circulating Water System (CWS), an important non-safety-related system that impacts the power generation capability of the plant site. Achieving riskinformed condition-based Predictive Maintenance (PdM) on the CWS will result in significant economic benefits, and the developed methodologies can also be applied to other plant systems. This approach supports an industry goal of ensuring that nuclear power generation remains a viable, economically competitive option in the energy market. Operation and Maintenance (O&M) costs include labor-intensive Preventive Maintenance (PM) programs that involve manually performed inspection, calibration, testing, and maintenance of plant assets at periodic frequencies as well as time-based replacement of assets, irrespective of condition. This project offers an alternative by focusing on riskinformed condition-based maintenance to reduce O&M costs while still maintaining plant health and safety. This report summarizes the progress made toward achieving a risk-informed condition-based maintenance approach. The research and development (R&D) activities presented in this report are associated with development of a nuclear digital platform application, integration of fault signature models, and automated work management processes. The fault signatures and Machine Learning (ML) models are key components in predictive analytics and are heavily leveraged to improve the insights received by existing plant process data sources. Availability of the analysis results within a centralized digital platform enhances efficiency by enabling automation of activities otherwise performed manually. Personnel are presented with enhanced information that can be used to evaluate plant status and risks. Utilizing the enhancements to data analytics supports automated responses, (i.e. issuance of work orders) to address developing equipment faults and thus preventing forced, unplanned shutdowns of components or systems. The R&D activities described within this report lay the foundation for developing and demonstrating a digital automated platform to centralize the implementation of condition monitoring and response to equipment faults. The digital automated platform is cloud-based and designed to enable improved efficiency of plant processes. The digital platform includes content related to maintenance optimization, fault signature analysis, and plant records, which can all be used to support efficiencies when located within a centralized digital platform. These efficiencies could be further enhanced when deployed through industry-wide deployment of the technology to improve insights and processes based upon economies of scale.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Interactive Exploration of High-Dimensional Phase Diagrams

High-dimensional thermodynamic phase stability databases are becoming increasingly common due to the convergence of three recent trends: (i) the widespread interest in so-called “high-entropy” alloys, (ii) the availability of high-throughput computational assessments of phase stability in broad composition spaces and (iii) the ongoing development of ever-increasingly broad, multicomponent, multiphase CALPHAD databases. Although automated computational tools can readily process such high-dimensional data, scientists are often unable to visualize the relevant phase relations, an ability that is crucial to gaining an intuitive understanding of the stability constraints governing materials design. The present work addresses this need by providing algorithms that enable the interactive exploration of phase equilibria in high-dimensional spaces. These algorithms concentrate the complex nonlinear nonsmooth optimization needed into a preprocessing step that generates a large number of high-dimensional yet elementary graphical primitives. Furthermore, these primitives can then be cross-sectioned to yield 3-dimensional views in a computationally efficient manner that enables an interactive exploration of high-dimensional spaces. All of these operations are highly parallelizable, thus facilitating scaling of this method to large data sets.

36 MATERIALS SCIENCE↗

Accelerating Advanced Light Source Science Through Multi-Facility HPC Workflows

Synchrotron light sources support a wide array of techniques to investigate materials, often producing complex, high-volume data that challenge traditional workflows. At the Advanced Light Source (ALS), we developed infrastructure to move microtomography data over ESnet to ALCF and NERSC, where CPU- and GPU-based algorithms generate 3D reconstructed volumes of experimental samples. We employ two data movement and reconstruction models: real-time processing as data streams directly to NERSC compute nodes, and automated file transfer to NERSC and ALCF file systems. The streaming pipeline provides users with feedback in under ten seconds, while the file-based workflow produces high-quality reconstructions suitable for deeper analysis in 20-30 minutes. This infrastructure enables users to utilize HPC resources without direct access to backend systems. We plan to extend this architecture to more endstations, supporting our beamline scientists and users.

Abramov, David↗

Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

Machine learning (ML)-accelerated discovery requires large amounts of high-fidelity data to reveal predictive structure–property relationships. For many properties of interest in materials discovery, the challenging nature and high cost of data generation has resulted in a data landscape that is both scarcely populated and of dubious quality. Data-driven techniques starting to overcome these limitations include the use of consensus across functionals in density functional theory, the development of new functionals or accelerated electronic structure theories, and the detection of where computationally demanding methods are most necessary. When properties cannot be reliably simulated, large experimental data sets can be used to train ML models. In the absence of manual curation, increasingly sophisticated natural language processing and automated image analysis are making it possible to learn structure–property relationships from the literature. Finally, models trained on these data sets will improve as they incorporate community feedback.

36 MATERIALS SCIENCE↗

Directional Laplacian Centrality for Cyber Situational Awareness

Cyber operations is drowning in diverse, high-volume, multi-source data. To get a full picture of current operations and identify malicious events and actors, analysts must see through data generated by a mix of human activity and benign automated processes. Although many monitoring and alert systems exist, they typically use signature-based detection methods. We introduce a general method rooted in spectral graph theory to discover patterns and anomalies without a priori knowledge of signatures. We derive and propose a new graph-theoretic centrality measure based on the derivative of the graph Laplacian matrix in the direction of a vertex. To build intuition about our measure, we show how it identifies the most central vertices in standard network datasets and compare to other graph centrality measures. Finally, we focus our attention on studying its effectiveness in identifying important IP addresses in network flow data. Using both real and synthetic network flow data, we conduct several experiments to test our measure’s sensitivity to two types of injected attack profiles and show that vertices participating in injected attack profiles exhibit noticeable changes in our centrality measures, even when the injected anomalies are relatively small, and in the presence of simulated network dynamics.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Using Systems Theoretic Process Analysis and Causal Analysis to Map and Manage Organizational Information to Enable Digitalization and Information Automation

The overarching goal of this Light Water Reactor Sustainability Program–supported research and development project is to provide planning tools and comprehensive guidance to utilities considering or undertaking full nuclear plant modernization. The results of this research will provide the nuclear industry with a comprehensive and usable solution, including guidance, lessons learned, methods, and planning tools. This research is currently working to provide guidance on digitalization and information automation to enable the evolution of data to information, insight, and action—thereby allowing utilities to operate safely and cost-competitively with all other electrical generation sources. Light Water Reactor Sustainability Program researchers have also recently started investigating how human and technology integration principles, information automation, and digitalization enable data evolution. These researchers are currently in the process of validating the use of System-Theoretic Process Analysis to define high-level safety constraints in the United States Nuclear Regulatory Commission’s problem identification and resolution process (i.e., a plant compliance information gathering activity). The next step in this research, which is described in the following sections of this report, is to map out data evolution in a use case to identify inefficiencies in another aspect of plant compliance information gathering and communication activities—event investigations and root cause analyses.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Automation of Laser Plasma Focused Ion Beam Microscopy for Next-Gen Energy Materials

Automation can revolutionize the use of ultrafast laser ablation and plasma-focused ion beam (PFIB) techniques for high-throughput, reproducible cross-sectioning and various sample preparation in materials characterization. As these methods become essential for analyzing complex energy materials and next-generation devices, efficient, standardized workflows are needed to minimize variability and enhance precision. This work highlights our advancements in developing automated processes for sample preparation that integrates machine learning, workflow optimization, and large-scale data acquisition to improve efficiency and scalability in applications such as electrolyzers, photovoltaic cells, and microelectronics. To streamline cross-sectioning and lamella fabrication, we have implemented fully automated workflows that standardize laser ablation and PFIB milling sequences. These workflows incorporate pre-programmed protocols for material removal, alignment, and thinning, reducing user intervention and ensuring consistency across different sample types. Machine learning algorithms further enhance automation by predicting optimal milling strategies and adapting parameters based on material properties and sectioning requirements. This approach significantly improves throughput while maintaining the structural integrity of prepared samples for high-resolution imaging and analysis, including transmission electron microscopy. Beyond sample preparation, our automation platform enables the acquisition of large, high-resolution datasets through serial sectioning, image alignment, and 3D reconstruction. These automated routines facilitate multi-scale characterization, capturing structural and compositional details from the nanoscale to the device level. By reducing variability and increasing efficiency, our automated approach enhances defect analysis, failure diagnostics, and process optimization, accelerating advancements in materials research and device engineering.

36 MATERIALS SCIENCE↗

FrESCO: Framework for Exploring Scalable Computational Oncology

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information from electronic health records. This data is then used to create summary statistics about cancer incidence and mortality to facilitate population health monitoring. Extracting phenotypic information from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating the information extraction process from cancer pathology reports has the potential to improve data quality by extracting information in a consistent manner across registries. It can also improve patient outcomes by reducing the time from diagnosis, enabling rapid case ascertainment for clinical trials. Here we present FrESCO, a modular deep-learning natural language processing (NLP) library initially designed for extracting pathology information from clinical text documents. This repository is not solely limited to clinical medical text, but may also be used by researchers just getting started with NLP methods and those looking for a robust solution for their classification problems.

60 APPLIED LIFE SCIENCES↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Towards operational atmospheric correction of airborne hyperspectral imaging spectroscopy: Algorithm evaluation, key parameter analysis, and machine learning emulators

Atmospheric correction of airborne hyperspectral imaging spectroscopy (AHIS) to obtain high-quality surface reflectance is the prerequisite for remote sensing applications. Over the last decades, different atmospheric correction methods have been developed based on radiative transfer models (RTMs), however, the relative performances of different algorithms are unclear. Automated operational atmospheric correction methods to process large-volume AHIS data in a high-accurate and high-throughput manner are still lacking. Therefore, this study proposed an operational atmospheric correction pipeline for deriving surface reflectance from AHIS data. To ensure the accuracy and efficiency of the pipeline, we focused on three specific aspects: (1) selecting a suitable RTM for the development of atmospheric lookup tables (LUTs) by comparing the commercial MODerate resolution atmospheric TRANsmission (MODTRAN) and open-sourced Library for Radiative TRANsfer (LibRadTRAN) models, where the widely-used software, Atmospheric/Topographic Correction for Airborne Imagery (ATCOR), was used as benchmarks; (2) identifying key atmospheric correction parameters and determining suitable sources for parameter retrievals including AHIS, Moderate Resolution Imaging Spectroradiometer (MODIS), and AErosol RObotic NETwork (AERONET); and (3) testing the performance of using machine learning emulators to speed up the RTM-based atmospheric correction. Results indicate that (1) atmospheric correction based on MODTRAN LUTs can produce surface reflectance accurately with mean absolute errors < 0.05 and cosine similarities > 0.98 compared to field measurements, which is comparable to the software ATCOR and slightly outperforms the LibRadTRAN LUTs; (2) sobol global sensitivity analysis demonstrates that in the atmospheric correction, visibility and water vapor are two key parameters that can be accurately derived from AHIS in contrast to MODIS or AERONET data; and (3) Random Forest emulators can produce accurate estimations of surface reflectance with mean absolute errors < 0.03 and cosine similarities > 0.98 for higher processing efficiency and determine a suitable set of wavelengths for retrieving atmospheric visibility and water vapor. In conclusion, the proposed atmospheric correction pipeline also improved the four-stream radiative transfer theory for airborne applications by considering adjacent effects from airborne surrounding pixels and can also be applied for atmospheric correction of hyperspectral data from spaceborne missions.

47 OTHER INSTRUMENTATION↗

recon3d

SAND2025-00533O recon3d is a software tool that provides automated 3D reconstruction and meshing capabilities. It processes labeled 3D image data from various sources, starting from image stacks, and calculates 3D feature distributions like size, shape, and location. The software also has tools for downscaling rectilinear grid data and creating tetrahedral meshes directly from image data. recon3d can be used by novice users via the command line with a properly formatted configuration file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Emery, John↗

The Integration and Mapping of an Open-Source National Well Resource to Inform Geologic Carbon Storage Site Selection and Risk Prevention: The CO2-Locate Database

Geologic carbon storage (GCS) offers a way to capture and permanently store CO₂ from fossil fuel operations in underground geologic structures, aiding in the transition to a carbon-neutral energy economy. However, CO₂ injection sites can experience gas leakage through existing wells that penetrate storage reservoirs, making knowledge of well locations and characteristics crucial for permitting, infrastructure reusability, and risk assessment in GCS. Currently, public wellbore data from state, federal, and tribal entities are inconsistent and fragmented, with gaps and redundancies. To address this, the National Energy Technology Laboratory (NETL) developed CO2-Locate, an open-source, geospatial database and online application. CO2-Locate integrates over 50 data sources from federal, state, and tribal entities, creating a standardized national well database. Funded by the Bipartisan Infrastructure Law, the database is publicly available through the Energy Data eXchange (EDX) and viewable via the CO2-Locate web mapping application. This tool allows users to query, filter, and visualize well data to support GCS planning, permitting, and risk assessments. This presentation covers the methods used to create CO2-Locate, including data acquisition, processing, attribute mapping, and integration, much of which is automated for future updates. The web mapping application and its role in GCS site selection will also be discussed.

Tetteh, Daniel A.↗

Pyrolysis Molecular Beam Mass Spectrometry_Analysis_of_Natural_Variants_of_Poplulus_Trichocarpa_Leaves

Select leaves from natural variants of Poplar (Populus Trichocarpa) grown in a greenhouse at Oak Ridge National Laboratory were analyzed by Pyrolysis-Molecular Beam Mass Spectrometry (Py-MBMS). Leaves were harvested, cryomilled and kept frozen until analysis. Py-MBMS analysis was conducted using approximately 4 mg of biomass and each sample was analyzed in duplicate. A Frontier PY2020 unit pyrolyzed samples at 500°C for 30 s in 80 µL deactivated stainless steel cups. An Extrel Super-Sonic MBMS Model Max 1000 was used to collect mass spectral data fromm/z30 to 450 at 17 eV and processed using Merlin Automation software (V3). Spectral ion intensities were normalized to the total ion chromatogram signal for each sample for analysis of spectral variance. Lignin content (wt %) was estimated based on relative responses from standards of known Klason lignin content using mean-normalized ion intensities ofm/z120, 124 (G), 137 (G), 138 (G), 150 (G), 152, 154 (S), 164 (G), 167 (S), 168 (S), 178 (G), 180, 181, 182 (S), 194 (S), 208 (S) and 210 (S) where G indicates guaiacyl-derived ions, S indicates syringyl-derived ions, and other ions either derive from other lignin monomers or multiple sources. Ratios of S and G lignin monomer units (S/G) were obtained by dividing the sum of S-based ions by the sum of G-based ions using mean-normalized ion intensities.

CBI↗

Pyrolysis_Molecular_Beam_Mass_Spectrometry_Analysis_of_Specific_Switchgrass_Genotypes

Select natural variant switchgrass genotypes grown in Tifton, GA were analyzed by Pyrolysis-Molecular Beam Mass Spectrometry (Py-MBMS). Biomass was harvested, milled, several genotypes were analyzed with and without being destarched and extracted with ethanol prior to analysis (indicated with -DE if destarched and extracted). Py-MBMS analysis was conducted using approximately 4 mg of biomass and each sample was analyzed in duplicate. A Frontier PY2020 unit pyrolyzed samples at 500°C for 30 s in 80 µL deactivated stainless steel cups. An Extrel Super-Sonic MBMS Model Max 1000 was used to collect mass spectral data fromm/z30 to 450 at 17 eV and processed using Merlin Automation software (V3). Spectral ion intensities were normalized to the total ion chromatogram signal for each sample for analysis of spectral variance. Lignin content (wt %) was estimated based on relative responses from standards of known Klason lignin content using mean-normalized ion intensities ofm/z120, 124 (G), 137 (G), 138 (G), 150 (G), 152, 154 (S), 164 (G), 167 (S), 168 (S), 178 (G), 180, 181, 182 (S), 194 (S), 208 (S) and 210 (S) where G indicates guaiacyl-derived ions, S indicates syringyl-derived ions, and other ions either derive from other lignin monomers or multiple sources. Ratios of S and G lignin monomer units (S/G) were obtained by dividing the sum of S-based ions by the sum of G-based ions using mean-normalized ion intensities.

CBI↗

Pyrolysis_Molecular_Beam_Mass_Spectrometry_Analysis_of_hybrid_cross_of_Populus_tremula_x_P_alba_717-1B4_and_overexpression_of_a_lectin_receptor-like_kinase_(PtLecRLK1)

Stem tissues from the hybrid poplarPopulus tremula × P. albaclone 717-1B4 and from lectin receptor-like kinase overexpression lines PP7 and PP19 were individually colonized with the ectomycorrhizal fungiLaccaria bicolorstrain S238N,Hyaloscypha finlandicastrain PMI746, orUmbelopsis vinaceastrain PMI3018, as well as with a mixed fungal inoculum; non-inoculated plants served as controls. Plants were grown in a greenhouse at Oak Ridge National Laboratory and harvested in January 2025. Stem samples were analyzed using Pyrolysis–Molecular Beam Mass Spectrometry (Py-MBMS). Stems were harvested, debarked, dried, milled, destarched and ethanol extracted prior to analysis. Py-MBMS analysis was conducted using approximately 4 mg of wood from biomass and each sample was analyzed in duplicate. A Frontier PY2020 unit pyrolyzed samples at 500°C for 30 s in 80 µL deactivated stainless steel cups. An Extrel Super-Sonic MBMS Model Max 1000 was used to collect mass spectral data fromm/z30 to 450 at 17 eV and processed using Merlin Automation software (V3). Spectral ion intensities were normalized to the total ion chromatogram signal for each sample for analysis of spectral variance. Lignin content (wt %) was estimated based on relative responses from standards of known Klason lignin content using mean-normalized ion intensities ofm/z120, 124 (G), 137 (G), 138 (G), 150 (G), 152, 154 (S), 164 (G), 167 (S), 168 (S), 178 (G), 180, 181, 182 (S), 194 (S), 208 (S) and 210 (S) where G indicates guaiacyl-derived ions, S indicates syringyl-derived ions, and other ions either derive from other lignin monomers or multiple sources. Ratios of S and G lignin monomer units (S/G) were obtained by dividing the sum of S-based ions by the sum of G-based ions using mean-normalized ion intensities.

CBI↗