Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Researcher Views on Water Data Accessibility, Usability, and Dissemination

Ultimately, the challenges that researchers face with finding, accessing, using, and disseminating water data affect their ability to do mission critical research for DOE and other sponsors. We interviewed PNNL researchers who support DOE’s Water Power Technologies Office and summarize their views on issues pertaining to data access, use, and dissemination. We also take a closer look at approaches to disseminating data to help inform WPTO and lab researcher decisions about incorporating data dissemination into future projects.

13 HYDRO ENERGY↗

PNNL Superfund Research Program Analytics (srpAnalytics Data Release v.1.0)

The OSU/PNNL superfund Research Program represents a longstanding collaboration to quantify Polycyclic Aromatic Hydrocarbons at various superfund sites in the Pacific Northwest and assess their potential impact on human health. To link the chemical measurements to biological activity, we describe the use of the zebrafish as a high-throughput developmental model of human exposure that provides quantitative measurements of the biological ramifications of exposure to toxicants. The PNNL Superfund have implemented a dose-response modelling pipeline to calculate benchmark dose parameters that enable the comparison of potency across chemicals and phenotypes. Our portal provides public access to this dataset and an interactive web site designed to enable exploration and re-use of this data by the scientific community at http://srp.pnnl.gov.

59 BASIC BIOLOGICAL SCIENCES↗

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram↗

Evaluation of Oak Ridge National Laboratory Health Physics Research Reactor Operation Data for Criticality Accident Alarm System Benchmark Creation

Over the last few decades, the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Working Group has focused on gathering critical and subcritical benchmark experiment data from different facilities around the world for inclusion in the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook). The compiled data can be used by criticality safety analysts to help validate their calculation tools and cross-section libraries for a variety of applications.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Discriminative Dimensionality Reduction using Deep Neural Networks for Clustering of LIGO Data

In this paper, leveraging the capabilities of neural networks for modeling the non-linearities that exist in the data, we propose several models that can project data into a low dimensional, discriminative, and smooth manifold. The proposed models can transfer knowledge from the domain of known classes to a new domain where the classes are unknown. A clustering algorithm is further applied in the new domain to find potentially new classes from the pool of unlabeled data. The research problem and data for this paper originated from the Gravity Spy project which is a side project of Advanced Laser Interferometer Gravitational-wave Observatory (LIGO). The LIGO project aims at detecting cosmic gravitational waves using huge detectors. However non-cosmic, non-Gaussian disturbances known as "glitches", show up in gravitational-wave data of LIGO. This is undesirable as it creates problems for the gravitational wave detection process. Gravity Spy aids in glitch identification with the purpose of understanding their origin. Since new types of glitches appear over time, one of the objective of Gravity Spy is to create new glitch classes. Towards this task, we offer a methodology in this paper to accomplish this.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Chemistry and Metallurgy Research Building Historical Perchlorate Data and Discussion

The Chemistry and Metallurgy Research Building (CMR), located at TA-03-0029, was built in 1952 and was occupied by Los Alamos National Laboratory (LANL) employees in 1953. The facility was built to perform actinide analytical chemistry and material characterization in support of LANL and Department of Energy (DOE) mission of stockpile stewardship and research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Shape-shifting Elephants: Multi-modal Transport for Integrated Research Infrastructure

Data Acquisition (DAQ) workloads form an important class of scientific network traffic that by its nature (1) flows across different research infrastructure, including remote instruments and supercomputer clusters, (2) has ever-increasing throughput demands, and (3) has ever-increasing integration demands---for example, observations at one instrument could trigger a reconfiguration of another instrument. Today's DAQ transfers rely on UDP and (heavily tuned) TCP, but this is driven by convenience rather than suitability. The mismatch between Internet transport protocols and scientific workloads becomes more stark with the steady increase in link capacities, data generation, and integration across research infrastructure.This position paper argues the importance of developing specialized transport protocols for DAQ workloads. It proposes a new transport feature for this kind of elephant flow: multi-modality involves the network actively configuring the transport protocol to change how DAQ flows are processed across different underlying networks that connect scientific research infrastructure. Multi-modality is a layering violation that is proposed as a pragmatic technique for DAQ transport protocol design. It takes advantage of programmable network hardware that is increasingly being deployed in scientific research infrastructure. The paper presents an initial evaluation through a pilot study that includes a Tofino2 switch and Alveo FPGA cards, and using data from a particle detector.

97 MATHEMATICS AND COMPUTING↗

The OCEAN ICE mooring compilation: a standardised, pan-Antarctic database of ocean hydrography and current time series

Continuous moored time series of temperature, salinity, pressure and current speed and direction are of great importance for understanding the continental shelf and under-ice-shelf dynamics and thermodynamics that govern water mass transformations and ice melting in and around Antarctic marginal seas. In these regions, icebergs and sea ice make ship-based mooring deployment and recovery challenging. Nevertheless, over decades, expeditions around the fringe of Antarctica sporadically deployed and recovered hundreds of moored instruments, including those facilitated through ice shelves boreholes. These datasets tend to be archived in a wide range of data centres, with, to our knowledge, no clear format standardisation. As a result, systematic analysis of historical mooring time series in the marginal seas is often challenging. Here we present the first version of a standardised pan-Antarctic moored hydrography and current time series compilation, with broad international contributions from data centres, research institutes and individual data owners. The mooring records in this compilation span over five decades, from the 1970s to the 2020s, providing an opportunity for a systematic study of the pan-Antarctic water mass transport and shelf connectivity. As a demonstration of the utility of this compilation, we present spectral analysis of the compiled current velocity time series, which unsurprisingly shows the dominating presence of tidal variability within most records. This component of the variability is fitted using multi-linear regression to tidal frequencies, and the tidal fit is removed from the original time series to leave de-tided variability. Given the limited record durations to months to years, de-tided variability is dominated by synoptic (3–10 d period), intraseasonal (10–80 d) and seasonal (∼6 months–1 year) signals. The spatial distribution of the kinetic energy integrated within frequency bands is presented and discussed within respective regional contexts, and future avenues of research are proposed. This data compilation is assembled under the endorsement of Ocean-Cryosphere Exchanges in ANtarctica: Impacts on Climate and the Earth System (OCEAN ICE) project (https://ocean-ice.eu/, last access: 23 October 2025) funded by the European Commission and UK Research and Innovation. It is available and regularly updated in NetCDF format with the SEANOE database at https://doi.org/10.17882/99922 (Zhou et al., 2024a).

54 ENVIRONMENTAL SCIENCES↗

Data-Driven PMU Noise Emulation Framework using Gradient-Penalty-Based Wasserstein GAN

Availability of phasor measurement unit (PMUs) data has led to research on data-driven algorithms for event monitoring, control and ensuring stability of the grid. Unavailability of infrequent critical event field PMU data with component failures is driving the need to generate realistic synthetic PMU data for research. The synthetic data from power system simulation software often neglect noise profiles of received phasors, thus creating some discrepancies between real PMU data and synthetic ones. To address this issue, this work presents an initial study on the noise characteristics of PMUs, as well as presenting models for recreating their unique noise signatures. The proposed method, utilizing the Wasserstein generative adversarial network with gradient penalty (WGAN-GP) architecture, provides an excellent benchmark for matching the noise distribution. One can use a well-learned GAN model to draw noise signatures from a distribution that seemingly mirrors the real PMU noise distribution, while also being able to be detached from the PMU data once the training is done. Based on the observed results and employed data-driven methodology, it is expected that the proposed methods can be adapted to replicate the behavior of other sensors, providing research and other applications with a tool for data synthesis and sensor characterization.

PMU↗

Data associated with Quaternary Research manuscript Latest Pleistocene glacial chronology and paleoclimate reconstruction for the East River watershed, Colorado, USA

The data here are associated with Quaternary Research manuscript Latest Pleistocene glacial chronology and paleoclimate reconstruction for the East River watershed, Colorado, USA and include data associated with cosmogenic exposure and depth profile dating as well as glacier-climate numerical modeling. Reconstructing Pleistocene glaciation timing and extent is vital for understanding paleoclimate. While late Pleistocene glaciation has been studied extensively in western North American mountain ranges, the glacial history of the western Elk Range in Colorado remains understudied, particularly in the East River watershed, a site of intense scientific focus. Here we use cosmogenic nuclide exposure and depth–profile dating methods to determine the timing of glaciation in the East River watershed. We use glacier modeling to reconstruct paleoglacier extents and quantify past climate conditions. Our findings indicate that the East River glacier retreated from its maximum position approximately 17–18 ka, moving to recessional positions between 13–15 ka, before experiencing more substantial retreat to high-elevation cirques around 12.9 ka. Glacier modeling suggests that the maximum ice extents at 17–18 ka could have been sustained by temperature depressions of approximately −6.5°C compared to modern conditions, assuming consistent precipitation. Additionally, the ice position at 13–15 ka could have been supported by temperature depressions of around −4.0°C. These results offer insights into the deglaciation timeline in the East River watershed and broader western Elk Range as well as paleoclimate conditions during the late Pleistocene, which may aid future research on critical zone evolution in the East River watershed.The data files include: 1. Table 1 In situ-produced 10Be sample data (.pdf, .csv, .xlsx)2. Table 2 In situ-produced 10Be exposure age results (.pdf, .csv, .xlsx)3. S1_A CRONUS calculator input assuming no erosion (.csv)4. S1_B CRONUS calculator input assuming erosion (.csv)5. S1_C CRONUS calculator results and comparison (.csv)6. S2_A In situ-produced 10Be sample data for depth profile (.csv)7. S2_B In situ-produced 10Be depth profile model input (.csv)8. S2_C In situ-produced 10Be depth profile model results (.csv)9. S3 Monthly cloudiness, rH, and windspeed data used in glacier climate model (.csv)10. S4 Glacier climate model results (.csv)

54 ENVIRONMENTAL SCIENCES↗

Lessons Learned from AskGDR: Usage and Impact Analysis of the Geothermal Data Repository's AI Research Assistant: Preprint

In October of 2024, the Department of Energy's (DOE) Geothermal Data Repository (GDR) team officially launched AskGDR, an AI research assistant resulting from the integration of a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets. AskGDR allows GDR users to ask deeper questions about the origin of datasets, the methods used to collect them, and the findings they help support. Using Retrieval Augmented Generation (RAG), AskGDR can be used to summarize findings spread across dozens of papers and technical reports or to extract relevant information describing a single data field. However, generative AI is experimental. The National Renewable Energy Laboratory (NREL) has been collecting metrics on AskGDR and documenting lessons learned during its deployment. This paper will outline the efficacy and impact of AskGDR through analysis of its use, operating costs, number and types of questions asked, and the quality of answers provided.

15 GEOTHERMAL ENERGY↗

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES↗

Optimizing Deep Learning Models for Climate-Related Natural Disaster Detection from UAV Images and Remote Sensing Data

This research study utilized artificial intelligence (AI) to detect natural disasters from aerial images. Flooding and desertification were two natural disasters taken into consideration. The Climate Change Dataset was created by compiling various open-access data sources. This dataset contains 6334 aerial images from UAV (unmanned aerial vehicles) images and satellite images. The Climate Change Dataset was then used to train Deep Learning (DL) models to identify natural disasters. Four different Machine Learning (ML) models were used: convolutional neural network (CNN), DenseNet201, VGG16, and ResNet50. These ML models were trained on our Climate Change Dataset so that their performance could be compared. DenseNet201 was chosen for optimization. All four ML models performed well. DenseNet201 and ResNet50 achieved the highest testing accuracies of 99.37% and 99.21%, respectively. This research project demonstrates the potential of AI to address environmental challenges, such as climate change-related natural disasters. This study’s approach is novel by creating a new dataset, optimizing an ML model, cross-validating, and presenting desertification as one of our natural disasters for DL detection. Three categories were used (Flooded, Desert, Neither). Our study relates to AI for Climate Change and Environmental Sustainability. Drone emergency response would be a practical application for our research project.

AI↗

Data to Accompany: Expanding the access of wearable silicone wristbands in community-engaged research through best practices in data analysis and integration

Wearable silicone wristbands are a rapidly growing exposure assessment technology that offer researchers the ability to study previously inaccessible cohorts and have the potential to provide a more comprehensive picture of chemical exposure within diverse communities. However, there are no established best practices for analyzing the data within a study or across multiple studies, thereby limiting impact and access of these data for larger meta-analyses. We utilize data from three studies, from over 600 wristbands worn by participants in New York City and Eugene, Oregon, to present a first-of-its-kind manuscript detailing wristband data properties. We further discuss and provide concrete examples of key areas and considerations in common statistical modeling methods where best practices must be established to enable meta-analyses and integration of data from multiple studies. Finally, we detail important and challenging aspects of machine learning, meta-analysis, and data integration that researchers will face in order to extend beyond the limited scope of individual studies focused on specific populations.

Bramer, Lisa M↗

Remote Instrumentation and Data Acquisition: An Internship Research Report

This report outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope’s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and lessons learned, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Fermilab]↗

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗