Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Temperature, Humidity, and Time-Lapse Video Data from the East River Watershed, Water Years 2024 and 2025

This dataset contains time-lapse imagery and distributed measurements of air temperature, relative humidity, dew point, and soil temperature across the East River basin from 3 October 2023 to 8 August 2025. Instruments were deployed at 19 sites as part of the DOE Grant: Seasonal Cycles Unravel Mysteries of Missing Mountain Water organized by Jessica Lundquist (University of Washington), Rosemary Carroll (Desert Research Institute), and Ethan Gutmann (National Center for Atmospheric Research). The data are published to support studies of surface climate or hydrologic processes in complex terrain. Measurements were collected with low-cost data loggers installed 2 m high on evergreen trees or buried just below the soil surface. Time-lapse cameras were deployed at three sites. Imagery from sites AP BONUS and AP5 (Avery Picnic) provides insight into large-scale seasonal snow cover variability. Imagery from site EL2 (Emerald Lake) shows smaller-scale snow patterns across a nearby meadow. Dataset files are organized by site and variable (air measurements, ground measurements, or time-lapse video). Air and ground measurements are packaged in LoggerData.zip, and time-lapse imagery is compiled into short videos stored in TimelapseVideos.zip. File-level metadata contains details for each file included in the dataset. A data dictionary provides units and descriptions for column or row names in all files. The locations metadata file describes site characteristics, locations, and associated GPS methods.

54 ENVIRONMENTAL SCIENCES↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Rapid Characterization and Statistical Analysis of High-Volume Field-Harvested Photovoltaic Connectors

Photovoltaic (PV) installations heavily depend on connectors for efficient module and string interconnections without requiring skilled labor. Yet this seemingly innocuous component of PV systems is a leading cause of module failures, multiple high-profile fires, and lawsuits in the PV industry. This work aims to answer critical questions regarding why connectors fail and the contributing factors to their failure. The study involves collecting and analyzing more than 17,000 field-harvested connectors from various solar installations across the United States. The vast dataset, which includes connector metadata, visual inspections, and resistance measurements, provides unprecedented insight into the state of health of PV connectors across the US, including the geographic locations, connector types, and installation practices most prone to failures. The work presented here describes a novel rapid characterization method for processing large numbers of connectors and is supported by parallel forensic analysis to discern the root causes of failures as well as a levelized cost of lifetime model to determine the economic ramifications of connector failure. Ultimately, the findings may inform PV developers about the best practices to extend connector longevity and lead to more resilient and reliable PV systems.

connectors↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access↗

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on proxy metrics such as bit operations (BOPs) that correlate poorly with hardware cost. This gap is particularly large for FPGA deployment, where cost is dominated by a multi-dimensional budget of lookup tables, DSPs, flip-flops, BRAM, and latency. We present the Surrogate Neural Architecture Codesign Package (SNAC-Pack), an open-source AutoML framework for hardware-aware neural architecture codesign and end-to-end FPGA deployment. SNAC-Pack runs a multi-objective global search with Optuna and NSGA-II, loading trials to a shared SQLite store that enables parallel workers across compute nodes. A hardware surrogate model outputs per-trial resource and latency estimates, avoiding the synthesis cost that would otherwise dominate the search loop. A local search stage then applies quantization-aware training (QAT) together with iterative magnitude pruning in a combined compression loop, after which the final model is synthesized to FPGA firmware via the hls4ml Python library. A YAML configuration and an optional agentic frontend let users run the pipeline on new datasets without modifying the framework. We demonstrate SNAC-Pack on jet classification at the Large Hadron Collider and superconducting qubit readout, discovering compact architectures that match or exceed strong baselines on the task metric while reducing FPGA resource utilization and, in the qubit readout case, reducing the design space exploration process from months of manual fine-tuning to hours of automated search.

Weitz, Jason [UC, San Diego]↗

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

Rapid characterization and failure analysis of 6276 rooftop-harvested photovoltaic connectors

Photovoltaic (PV) connectors, which link modules in series and connect PV strings in parallel, have increasingly been recognized as a primary contributor to PV system failures and a source of numerous fire incidents. However, publicly available data on the rates and types of connector failures are scarce, primarily due to the proprietary nature of the information and the need for comprehensive analysis. This study represents the first large-scale investigation of harvested PV connectors, drawing from a dataset of 6276 connectors from residential rooftop solar systems across the United States. The outcome of this work is twofold: 1) we have established a rapid characterization method for large populations of harvested connectors, incorporating visual inspection, resistance measurements, and X-ray imaging; and 2) the analysis made possible by our rapid-processing method has revealed, for a population of connector models provided by a single rooftop installer, failure statistics and insights for various connector makes and models, installation practices, operating currents, and internal component displacements. This research identifies common failure modes that could be considered in future connector designs standards, and operations and maintenance practices, to ultimately improve the reliability of this vital component of PV infrastructure.

MC4↗

The Role of the U.S. Electric Distribution System in Serving Data Center and Other Large Loads

The rapid expansion of data centers in the United States is reshaping how the electric distribution system must plan for and accommodate large load interconnections. This report evaluates the role of the distribution grid in serving these loads, from small edge facilities to hyperscale campuses. Using national datasets, utility filings, and industry studies, we assess demand growth, reliability requirements, interconnection thresholds, and infrastructure needs at substations and feeders. The analysis highlights the mismatch between fast data center development timelines and slower utility planning and construction cycles, as well as strategies such as phased energization, on-site generation, hosting capacity maps, and structured interconnection frameworks. While focused on data centers, the insights also apply to other high-density loads such as advanced manufacturing, hydrogen production, and electrified transportation. The report concludes with approaches to align planning processes, transparency tools, and regulatory frameworks so utilities can manage new large loads in ways that support a reliable and resilient grid.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications

Cybersecurity breaches targeting electrical substations constitute a significant threat to the integrity of the power grid, necessitating comprehensive defense and mitigation strategies. Any anomaly in information and communication technology (ICT) should be detected for secure communications between devices in digital substations. This paper proposes large language models (LLMs), e.g., ChatGPT, for the cybersecurity of IEC 61850-based communications. Multi-cast messages such as generic object oriented system events (GOOSE) and sampled values (SV) are used for case studies. The proposed LLM-based cybersecurity framework includes, for the first time, data pre-processing of communication systems and human-in-the-loop (HITL) training (considering the cybersecurity guidelines recommended by humans). The results show a comparative analysis of detected anomaly data carried out based on the performance evaluation metrics for different LLMs. A hardware-in-the-loop (HIL) testbed is used to generate and extract a dataset of IEC 61850 communications.

ChatGPT↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives. This paper will outline the development, integration, output, and efficacy of the AskGDR LLM, including adherence to scientific rigor through improvements designed to increase the accuracy of generated answers, avoid speculation, and provide proper references for all resources used.

access↗

Accuracy versus precision in boosted top tagging with the ATLAS detector

The identification of top quark decays where the top quark has a large momentum transverse to the beam axis, known as top tagging , is a crucial component in many measurements of Standard Model processes and searches for beyond the Standard Model physics at the Large Hadron Collider. Machine learning techniques have improved the performance of top tagging algorithms, but the size of the systematic uncertainties for all proposed algorithms has not been systematically studied. This paper presents the performance of several machine learning based top tagging algorithms on a dataset constructed from simulated proton-proton collision events measured with the ATLAS detector at $\sqrt{s}$ = 13 TeV. The systematic uncertainties associated with these algorithms are estimated through an approximate procedure that is not meant to be used in a physics analysis, but is appropriate for the level of precision required for this study. The most performant algorithms are found to have the largest uncertainties, motivating the development of methods to reduce these uncertainties without compromising performance. To enable such efforts in the wider scientific community, the datasets used in this paper are made publicly available.

47 OTHER INSTRUMENTATION↗

A Framework for Compressing Unstructured Scientific Data via Serialization

We present a general framework for compressing unstructured scientific data with known local connectivity. A common application is simulation data defined on arbitrary finite element meshes. The framework employs a greedy topology preserving reordering of original nodes which allows for seamless integration into existing data processing pipelines. This reordering process depends solely on mesh connectivity and can be performed offline for optimal efficiency. However, the algorithm’s greedy nature also supports on-the-fly implementation. The proposed method is compatible with any compression algorithm that leverages spatial correlations within the data. The effectiveness of this approach is demonstrated on a large-scale real dataset using several compression methods, including MGARD, SZ, and ZFP.

Reshniak, Viktor [ORNL] (ORCID:0000000315454462)↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

PV Degradation Modeling: Applying Geospatial Workflows with "PVDeg"

Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.

14 SOLAR ENERGY↗

Finding Nuclear Clusters in the Short-Baseline Near Detector

Production of nuclear clusters such as deuterons, tritons, helions and alpha particles has only recently been simulated for neutrino-nucleus interactions, and has not yet been measured in neutrino experiments. The Short-Baseline Near Detector (SBND) is the first Liquid Argon Time Projection Chamber (LArTPC) with high enough resolution and a large enough neutrino flux to accurately measure the cross-section for the production of heavier-than-proton fragments in neutrino interactions. SBND is a 112-ton LArTPC, and lies 110 m downstream from the Booster Neutrino Beam target, where it has already collected the world's largest dataset of neutrino-argon interactions. In these processes, the intranuclear cascade and nuclear de-excitation stages can both produce nuclear clusters. With LArTPC technology's excellent reconstruction capabilities and SBND's unprecedented neutrino statistics, a measurement of these nuclear clusters could distinguish between nuclear models and improve neutrino energy reconstruction. This poster presents the status of simulation, reconstruction, and prospects for a measurement of nuclear clusters in SBND.

Beever, Anna [Sheffield U.] (ORCID:000900069339924↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗