Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Expert Elicitation on Wind Farm Control

Wind farm control is an active and growing field of research in which the control actions of individual turbines in a farm are coordinated, accounting for inter-turbine aerodynamic interaction, to improve the overall performance of the wind farm and to reduce costs. The primary objectives of wind farm control include increasing power production, reducing turbine loads, and providing electricity grid support services. Additional objectives include improving reliability or reducing external impacts to the environment and communities. In 2019, a European research project (FarmConners) was started with the main goal of providing an overview of the state-of-the-art in wind farm control, identifying consensus of research findings, data sets, and best practices, providing a summary of the main research challenges, and establishing a roadmap on how to address these challenges. Complementary to the FarmConners project, an IEA Wind Topical Expert Meeting (TEM) and two rounds of surveys among experts were performed. From these events we can clearly identify an interest in more public validation campaigns. Additionally, a deeper understanding of the mechanical loads and the uncertainties concerning the effectiveness of wind farm control are considered two major research gaps.

17 WIND ENERGY↗

ggtaxplot v 0.0.1

ggtaxplot is an R package designed to process and visualize taxonomic data through a taxonomic river plot. This package is ideal for researchers and data scientists who need to visualize taxonomic data. ggtaxplot function processes data and generates a taxonomic river plot, allowing users to visualize the distribution of taxa across different samples.

Coclet, Clement [Lawrence Berkeley National Labora↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Electrification Analysis: Container Ports' Cargo Handling Equipment

This one-page highlight details the key takeaways from a project that utilized NREL's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline, the Electrification Analysis of Container Ports' Cargo Handling Equipment project. This project created a scalable solution to model energy demand per shipping container moved (kWh/TEU) for an all-electric cargo handling equipment fleet located at a maritime port. The model allows stakeholders to understand energy demand at each electric vehicle (EV) equipment level and is easily scalable to container demand and EV adoption rate projections.

ADVANCED PROPULSION SYSTEMS↗

Mathematics: The Tao of Data Science

The two pieces, "Ten Research Challenge Areas in Data Science" by Jeannette M. Wing and “Challenges and Opportunities in Statistics and Data Science: Ten Research Areas” by Xuming He and Xihong Lin, provide an impressively complete list of data science challenges from luminaries in the field of data science. They have done an extraordinary job, so this response offers a complementary viewpoint from a mathematical perspective and evangelizes advanced mathematics as a key tool for meeting the challenges they have laid out. Notably, we pick up the themes of scientific understanding of machine learning and deep learning, computational considerations such as cloud computing and scalability, balancing computational and statistical considerations, and inference with limited data. We propose that mathematics is an important key to establishing rigor in the field of data science and as such has an essential role to play in its future.

97 MATHEMATICS AND COMPUTING↗

Data From: Influence of Hydrological Perturbations and Riverbed Sediment Characteristics on Hyporheic Zone Respiration of CO2 and N-2, Journal of Geophysical Research-Biogeosciences

This data package contains pumping data (.txt), parameter matrices, and R code (.R, .RData) to perform bootstrapping for parameter selection for the bioclogging model development. The pumping data were collected from the Russian River Riverbank Filtration site located in Sonoma County, California from 2010-2017 from three riverbank collection wells located alongside the study site. The pumping data is directly correlated with water table oscillations, so the code performs these correlations and simulates stochastic versions of water table oscillations. See Metadata Description.pdf for full details on dataset production. This dataset must be used with the R programming language. This dataset and R code is associated with the publication "Influence of Hydrological Perturbations and Riverbed Sediment Characteristics on Hyporheic Zone Respiration of CO2 and N-2"This research was supported by the Jane Lewis Fellowship from the University of California, Berkeley, the Sonoma County Water Agency (SCWA), the Roy G. Post Foundation Scholarship, the U.S. Department of Energy, Office of Science Graduate Student Research (SCGSR) Program, U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research under award DE-AC02-05CH11231, and the UFZ-Helmholtz Centre for Environmental Research, Leipzig, Germany.

54 ENVIRONMENTAL SCIENCES↗

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES↗

Deep Learning-Based Failure Prognostic Model for PV Inverter Using Field Measurements

Here, this study presents a novel approach for the precise monitoring and prognosis of photovoltaic (PV) inverter status, which is crucial for the proactive maintenance of PV systems. It addresses the gaps in traditional model-based methods, which tend to neglect the overall reliability of inverters, and the limitations of data-driven approaches that largely depend on simulated data. This research presents a robust solution applicable to real-world scenarios. The proposed data-driven model for PV inverter failure prognosis employs actual inverter measurements, integrating various operational and weather-related factors based on domain knowledge. This approach effectively represents inverter stressors and operational status. Utilizing an Enhanced Siamese Convolutional Neural Network (ESCNN), the model merges operational data with domain knowledge features, redefining the prognosis challenge as a classification task. Furthermore, the paper discusses an ESCNN-based real-time inverter failure monitoring method developed on the well-trained model. The proposed models are rigorously trained and tested with real inverter data and a novel filtering method is included to address accidental failures in practical scenarios. The results validate the model's efficacy, and the directions for future research are also outlined.

42 ENGINEERING↗

Quantum Computing – Real-Time Data Processing from a Dilution Refrigerator

This poster presents a real-time data visualization system for monitoring and displaying data from a Bluefors control unit connected to a dilution refrigerator. The goal is to provide an aesthetically pleasing and user-friendly web-based interface that continuously updates and displays critical data such as temperature and flow rates at various points within the refrigerator. Leveraging WebSocket technology, the system establishes multiple connections to the dilution refrigerators, enabling simultaneous monitoring of several units. The interactive webpage dynamically updates the data, providing researchers and operators with instant insights into the system's performance. The system's ability to stream and visualize data in real-time enhances the understanding of the dilution refrigerator's behavior, aiding in optimizing its operation and ensuring efficient scientific experiments. The application's web-based nature makes it easily accessible from any device with internet connectivity, promoting seamless collaboration and remote monitoring capabilities. Overall, this poster offers a comprehensive solution for real-time data visualization and analysis of dilution refrigerators, catering to the needs of researchers and scientists working in low-temperature physics and quantum computing fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Making Common Fund data more findable: catalyzing a data ecosystem

The Common Fund Data Ecosystem (CFDE) has created a flexible system of data federation that enables researchers to discover datasets from across the US National Institutes of Health Common Fund without requiring that data owners move, reformat, or rehost those data. This system is centered on a catalog that integrates detailed descriptions of biomedical datasets from individual Common Fund Programs’ Data Coordination Centers (DCCs) into a uniform metadata model that can then be indexed and searched from a centralized portal. This Crosscut Metadata Model (C2M2) supports the wide variety of data types and metadata terms used by individual DCCs and can readily describe nearly all forms of biomedical research data. We detail its use to ingest and index data from 11 DCCs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Extracting Material Property Measurements from Scientific Literature with Limited Annotations

Extracting material property data from scientific text is pivotal for advancing data-driven research in chemistry and materials science; however, the extensive annotation effort required to produce training data for named entity recognition (NER) models for this task often makes it a barrier to extracting specialized data sets. Here, in this work, we present a comparative study of the conventional, supervised NER methodology to alternative few-shot learning architectures and large language model (LLM)-based approaches that mitigate the need to label large training data sets. We find that the best-performing LLM (GPT-4o) not only excels in directly extracting relevant material properties based on limited examples but also enhances supervised learning through data augmentation. We supplement our findings with error and data quality assessments to provide a nuanced understanding of factors that impact property measurement extraction.

36 MATERIALS SCIENCE↗

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit↗

A 30-yr high-resolution weather research and forecasting model downscaling data over California and Nevada

This dataset presents a 30-year high resolution meteorological dataset obtained using the WRF model (Advanced version Research WRF version 4.4). We used WRF and European Centre for Medium-Range Weather Forecasts Reanalysis v5 as initial and boundary conditions to generate gridded meteorological variables. A large number of surface weather stations was used for model validation. A multi-physics analysis was first developed to identify a good physics suite extended from 6 November 00 UTC to 10 November 23 UTC, 2018, which included the Camp Fire in northern California. Based on the best physics suite, the downscaling dataset extends from 1 December to 28 February, 1990–2021 and the horizontal domain has 1.5 km grid spacing covering the entire states of California and Nevada in the United States. Comparisons between hourly surface observations and WRF simulations of air temperature, relative humidity and wind speeds show mean absolute errors on the order of (1.6-2.0 C), (10 %) and 1.2–1.5 m s -1 , respectively.

54 ENVIRONMENTAL SCIENCES↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Quality Assurance Program Plan for SFR Metallic Fuel Data Qualification

This document contains an evaluation of the applicability of the current Quality Assurance Standards from the American Society of Mechanical Engineers Standard NQA-1 (NQA-1) criteria and identifies and describes the quality assurance process(es) by which attributes of historical, analytical, and other data associated with sodium-cooled fast reactor [SFR] metallic fuel will be evaluated. This process is being instituted to facilitate validation of data to the extent that such data may be used to support future licensing efforts associated with advanced reactor designs. The initial data to be evaluated under this program were generated during the US Integral Fast Reactor program between 1984-1994, where the data include, but are not limited to, research and development data and associated documents, test plans and associated protocols, operations and test data, technical reports, and information associated with past United States Nuclear Regulatory Commission reviews of SFR designs. It is recognized that managing the data generated by large research and development projects presents a significant challenge for retaining data integrity and availability. American Society of Mechanical Engineers Standard NQA-1 (NQA-1) 2008/2009a provides appropriate requirements for this plan.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Developing new pathways for energy and environmental decision-making in India: a review

Abstract India faces a dual challenge of economic development and responding to climate change. Although India’s per capita emissions are well below global average, the country is one of the world’s largest greenhouse gas emitters. Indian policymakers and stakeholders require high-quality data and research to assess low-emissions, sustainable development strategies. Peer-reviewed literature is a key source of this information and also a key venue for conversation amongst research leaders. This paper examines the recent peer-reviewed literature on India’s 2030 and 2050 pathways. We conducted a systematic literature review to identify key quantitative national modeling studies. From the 34 studies identified, we synthesized scenario data to draw common conclusions and identify critical research gaps. The main focus was on examining the coverage and the state of information available on low-carbon pathways. Overall, we find a few scenarios that are potentially consistent with a 2070 net-zero goal, but more limited assessment of pathways to reach net-zero emissions before this date. Mitigation pathways with greater ambition are required across all energy sectors to ensure a smooth transition to net-zero emissions by or before 2070. The scenarios confirm that reducing emissions to below 2 GtCO 2 yr −1 by mid-century would necessitate significant transformations of the Indian energy sector, such as, a decrease in unabated coal power capacity, transportation modal shift, and industrial process switching. The assessment also finds substantial differences in final energy estimates reported across studies, particularly in transportation. The lack of consistency in, and transparency about underlying drivers, assumptions, and even outputs across studies points to the critical need for the sorts of coordinated, multi-model studies that have proven exceptionally valuable for decision makers in other major emitting countries.

54 ENVIRONMENTAL SCIENCES↗