Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Offshore Geologic Carbon Storage Data Collection and Data Gaps Analysis

This is a TRS documenting the Offshore Geologic Carbon Storage Data Collection. It describes the Data Collection web application and its creation as well as an accompanying Data Gaps Assessment. We present an interactive data collection and data gaps analysis to aggregate, understand, and disseminate the data that are publicly available to support offshore GCS in the United States. This data collection and data gaps analysis can be leveraged by stakeholders to understand where GCS may be viable offshore, create GCS project analogs, and address challenges to GCS in offshore environments.

58 GEOSCIENCES

Comparative Economic Analysis Between Bioenergy and Forage Types of Switchgrass for Sustainable Biofuel Feedstock Production: A Data Envelopment Analysis and Cost–Benefit Analysis Approach

ABSTRACT The capacity to produce switchgrass efficiently and cost‐effectively across diverse environments can be pivotal in achieving the short‐ and medium‐term Sustainable Aviation Fuel targets set by the U.S. Department of Energy. This study evaluated the economic performance of forage‐ and bioenergy‐type switchgrass cultivars and their response to N fertilization under diverse marginal environments across the US Midwest that included Illinois (IL), Iowa (IA), Nebraska (NE), and South Dakota (SD). Data Envelopment Analysis (DEA) was used to evaluate the efficiency of 23 Decision‐Making Units (DMUs)—cultivar types and N fertilization rate combinations—while a cost–benefit analysis calculated their profitability over 5 years. Results showed that two energy‐type cultivars—“Independence” and “Liberty”—were superior economically to the forage cultivars. Independence performed best with the highest profit margin when fertilized at 56 kg N ha −1 , particularly in the US hardiness zone 6a (Urbana, IL). Liberty exhibited the highest profit margins in hardiness zone 5b (Madrid, IA, and Ithaca, NE) at 56 kg N ha −1 and showed exceptional profitability with 28 kg N ha −1 in hardiness zone 6b (Brighton, IL). Switchgrass cultivar “Carthage” showed better efficiency score and profitability results in hardiness zone 4b (South Shore, SD) at 56 kg N ha −1 . The profit trends observed in current study sites may indicate broader patterns across similar US hardiness zones. This study provides valuable insights for decision‐makers to optimize input strategies for biomass production of bioenergy switchgrass to meet renewable energy demands.

Arshad, Muhammad Umer [Department of Crop Sciences

Data Centers Gap Analysis [Slides]

Data centers and other large loads are a significant driver of unprecedented, near-term demand growth in the United States. Power system planners, utilities, regulators, and other stakeholders are grappling with how to integrate data centers on the system without comprising reliability, resiliency, and energy affordability. NLR is pursuing work to develop a siting and decision-making tool that would draw on power systems modeling expertise to achieve granular representation of trade-offs involved in data center sitting and development. This slide deck supports the same workstream by reviewing the literature to identify mitigation options to facilitate near-term integration of large loads and by presenting options for pursuing data development and/or modeling projects to improve representation of siting options.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Data-driven analysis of dipole strength functions using artificial neural networks

Here, we present a data-driven analysis of dipole strength functions across the nuclear chart, employing an artificial neural network to model nuclear dipole responses. We train the network on a dataset of experimentally measured dipole strength functions for 216 different nuclei. To assess its predictive capability, we test the trained model on an additional set of 10 new nuclei, where experimental data exist. We demonstrate that the artificial neural network not only accurately reproduces known data but also identifies potential inconsistencies in experimental datasets, indicating which results may warrant further review or possible rejection. For nuclei where experimental data are sparse or unavailable, the network confirms theoretical calculations, reinforcing its utility as a predictive tool in nuclear physics. Finally, utilizing the predicted electric dipole polarizability, we extract the value of the symmetry energy at saturation density and find it consistent with results from the literature.

artificial neural networks

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Benchmarking universal machine learning interatomic potentials for rapid analysis of inelastic neutron scattering data

The accurate calculation of phonons and vibrational spectra remains a significant challenge, requiring highly precise evaluations of interatomic forces. Traditional methods based on the quantum description of the electronic structure, while widely used, are computationally expensive and demand substantial expertise. Emerging universal machine learning interatomic potentials (uMLIPs) offer a transformative alternative by employing pre-trained neural network surrogates to predict interatomic forces directly from atomic coordinates. This approach dramatically reduces computation time and minimizes the need for technical knowledge. In this paper, we produce a phonon database comprising nearly 5000 inorganic crystals to benchmark the performance of several leading uMLIPs. We further assess these models in real-world applications by using them to analyze experimental inelastic neutron scattering data collected on a variety of materials. Through detailed comparisons, we identify the strengths and limitations of these uMLIPs, providing insights into their accuracy and suitability for fast calculations of phonons and related properties, as well as the potential for real-time interpretation of neutron scattering spectra. Our findings highlight how the rapid advancement of AI in science is revolutionizing experimental research and data analysis.

inelastic neutron scattering

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit

STM/S Grid LDOS Data and Analysis Code for Deciphering Majorana Zero Modes in Topological Superconductor

This dataset provides raw millikelvin scanning tunneling microscopy/spectroscopy (STM/S) grid spectroscopy data and Python analysis scripts supporting the manuscript “Deciphering Majorana Zero Modes in Topological Superconductor FeTe0.55Se0.45 with Machine-Learning-Assisted Spectral Deconvolution.” The dataset includes a raw grid spectroscopy file acquired on FeTe0.55Se0.45 at 40 mK under magnetic field, together with Python/Jupytext analysis scripts used for STM/S data processing, visualization, spectral deconvolution, Lorentzian peak fitting, feature extraction, machine-learning-assisted clustering, and figure generation. These files support the analysis of vortex-core local density of states and the identification of zero-bias-peak-related spectral components from complex in-gap states. The dataset is intended to provide a citable archival record of the data and analysis code associated with the published manuscript and to support transparency and reproducibility of the reported STM/S and machine-learning workflow.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES

SEAFORML (Smart Exploration and Analysis For Optimal and Robust Machine Learning)

The poster discusses data analysis of the WAVgraph database and applied machine learning methods for it. The database is a long-term project that seeks to be a comprehensive repository of information on cyber threats and is updated regularly. It was previously unanalyzed and unexplored. The goal was to learn more about it and its contents in order to have a better understanding and enable better use. The data analysis and discovery enabled further exploration through natural language processing, similarity, and clustering methods. The poster shows some of the insights from the analysis and explains the methods used for the machine learning applications.

24 - POWER TRANSMISSION AND DISTRIBUTION

Searches for New Long-Lived Particles and Upgrade to the ATLAS Inner Detector (Final Technical Report)

The search for new fundamental particles is one of the defining goals of the Large Hadron Collider (LHC). The discovery of the Higgs Boson by the ATLAS and CMS collaborations provided the capstone of the Standard Model of particle physics, but outstanding questions remain. Why does the Higgs boson have a mass of 125 GeV when its natural mass would be many orders of magnitude larger? Is there a universal symmetry which unites all three forces described by the Standard Model? Can that symmetry be extended to include gravity? Is dark matter, evidenced by astronomical observations, made of a particle that interacts via Standard Model forces with the rest of matter? Together, these motivations provide compelling arguments that new physical processes await discovery. This project addressed some outstanding questions about the fundamental particles and their interactions with the ATLAS experiment at the Large Hadron Collider. In particular, the project improved the discovery potential for new, long- lived particles produced via electroweak processes in proton-proton collisions and set world-leading limits on their existence for certain values of their potential mass and lifetime. To achieve this, the project developed new data analysis methods, developed new triggers to select events with new long-lived particles during data-taking of the ATLAS experiment, and analyzed the largest proton–proton collision dataset ever produced. The project also supported significant development of the data acquisition software for the upgrade to the ATLAS inner detector, the Inner TracKer (ITk). The upgrade of the ATLAS inner detector is essential to the success of the entire Phase II physics program on ATLAS. Personnel supported by the project provided support for integration, assembly, and testing of the inner two layers of the ITk pixel system during its prototype and pre-production phase. Four PhD students and two post-doctoral scholars were supported by the grant and received invaluable scientific training as part of the research endeavor. The students and postdocs gained essential professional skills in the areas of advanced data analysis techniques, statistical analysis of data and simulation, programming in C++ and Python, hardware and instrumentation development, and presentation and collaboration skills. Additionally, approximately ten undergraduate students supported through other funding sources participated in research activities synergistic with the goals of this project, receiving essential mentorship from the personnel supported by this project.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Life Cycle Inventories and Data Gap Analysis for Rare Earth Elements: Neodymium and Dysprosium from Mining to Magnets

The United States demand for Neodymium-Iron-Boron (NdFeB) magnets, produced from rare earth elements (REEs) such as (Nd) and Dysprosium (Dy), far exceeds its nascent domestic production capacity, rendering it reliant on vulnerable global supply chains dominated by China. To guide research and development investments in securing U.S. REE supply, defensible benchmark metrics across environmental, economic, and social dimensions are needed. In this study, we built globally-representative, process-based cradle-to-cradle life cycle inventories for Nd and Dy in NdFeB magnets lifecycles, encompassing primary material acquisition, beneficiation, smelting and refining, metal processing, specialty alloy and chemical transformation, subcomponent manufacturing, consumer application (use phase) and end-of-life management. We carried out detailed literature review, and applied process engineering principles to build industry-representative upscaled life cycle inventories for both metals. We used these models to conduct bottom-up literature review and gap analysis on existing literature, compilation of data sources for each life cycle stage (and transformations where necessary), and a preliminary technoeconomic analysis (TEA)/life cycle costing analysis (LCCA). Findings from this work emphasize the need for metal specific, representative REE LCIs to establish robust benchmarks for advancing sustainable REE technologies and guiding R&D in REE supply chains.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Automated Algorithms for Screening Electronic Parts for Aging using Power Spectra Analysis (PSA) Data

Understanding the age of semiconductor parts being built into devices and systems is of interest for manufacturing quality control. Power spectrum analysis (PSA) is a fast, non-destructive, sensitive method for examining semiconductor parts. This talk will cover the use of multivariate analysis on both PSA data and conventional current-voltage data generated prior to PSA analysis to create algorithms that can be automated to screen semiconductor parts for aging.

Multari, Rosalie A