Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Towards Auto-Generated Data Systems

After decades of progress, database management systems (DBMSs) are now the backbones of many data applications that we interact with on a daily basis. Yet, with the emergence of new data types and hardware, building and optimizing new data systems remain as difficult as the heyday of relational databases. In this paper, we summarize our work towards automating the building and optimization of data systems. Drawing from our own experience, we further argue that any automation technique must address three aspects: user specification, code generation, and result validation. We conclude by discussing a case study using videos data processing, along with opportunities for future research towards designing data systems that are automatically generated.

Computer Science↗

Reviews and syntheses: The promise of big diverse soil data, moving current practices towards future potential

Abstract. In the age of big data, soil data are more available and richer than ever, but – outside of a few large soil survey resources – they remain largely unusable for informing soil management and understanding Earth system processes beyond the original study. Data science has promised a fully reusable research pipeline where data from past studies are used to contextualize new findings and reanalyzed for new insight. Yet synthesis projects encounter challenges at all steps of the data reuse pipeline, including unavailable data, labor-intensive transcription of datasets, incomplete metadata, and a lack of communication between collaborators. Here, using insights from a diversity of soil, data, and climate scientists, we summarize current practices in soil data synthesis across all stages of database creation: availability, input, harmonization, curation, and publication. We then suggest new soil-focused semantic tools to improve existing data pipelines, such as ontologies, vocabulary lists, and community practices. Our goal is to provide the soil data community with an overview of current practices in soil data and where we need to go to fully leverage big data to solve soil problems in the next century.

54 ENVIRONMENTAL SCIENCES↗

Energy Systems Integration Facility (ESIF): World-Class Systems Integration Capabilities and Research

The Energy Systems Integration Facility (ESIF), located at the National Renewable Energy Laboratory (NREL) South Table Mountain campus, is a world-renowned user facility for research and development of modern, advanced, and clean energy technologies. ESIF is distinguished by its continuously evolving, highly integrated systems that span throughout the building, connecting research capabilities across multiple laboratories and test areas. The primary ESIF research systems include: [1] data, cyber, and control networks, [2] research electrical distribution buses (REDB), [3] thermal integration infrastructure, and [4] hydrogen systems. The data, cyber, and control networks provide monitoring, control, communication, automation, visualization, and time series data storage and tagging capabilities for research projects and ESIF systems, including facility safety functions. The REDB system consists of four dedicated AC and DC electrical power networks that can connect devices located across the facility through versatile, automatic circuit configuration to support complex power electronics experiments up to the megawatt-scale. The thermal integration infrastructure consists of three temperature-conditioned water loops that provide heating and cooling interfaces and capabilities for thermal energy research. The hydrogen systems provide megawatt-scale hydrogen production, drying, compression, high-pressure storage, and delivery to laboratory end uses, including hydrogen fuel cell vehicle fueling. The ESIF research systems interconnect and extend throughout the various lab areas of the facility to create elaborate networks composed of diverse technologies for cutting-edge research. The ESIF capabilities are operated and stewarded by the ESIF Research Operations group, who also actively upgrade and advance the systems to ensure they remain ahead of anticipated research - enabling the success of many pioneering energy integration projects. The poster, created by members of the ESIF Research Operations team, highlights and summarizes the four core integrated systems at ESIF. The poster was first presented at the internal NREL Energize Forum on May 13th, 2024, and received the "Best Poster" award.

capabilities↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

A systematic review of machine learning in groundwater monitoring

With increasing concerns about water scarcity, groundwater has become crucial since this resource provides most of the freshwater needs. However, various human and natural activities often contaminate the groundwater, making it unsuitable for use. Over the years, scientists and engineers have used many methods to predict and track groundwater contamination as part of environmental monitoring. Consequently, there is an urgent need for improved methods, particularly in the face of increasing contamination. Machine learning has sometimes been used to monitor groundwater, air quality, and climate. Traditional methods must be improved due to the complexity and large amount of environmental data. This includes using hybrid models that combine traditional and new techniques. Despite the use of machine learning in many scientific areas, there is a lack of comprehensive reviews focusing on its use in environmental monitoring, especially groundwater monitoring. We aim to fill this gap by exploring machine-learning applications in groundwater monitoring. We discuss relevant methods, their limitations, and future potential. We summarize research on automating data processing and model training using groundwater sensor data. Our research underscores the transformative potential of machine learning to revolutionize long-term groundwater monitoring and contamination detection, providing valuable insights for future research and practical applications.

AI/ML↗

RTe3_data

The purpose of this dataset is to provide data availability for a research paper. It contains experimental data used to analyze and present the results discussed in the research paper.

Physics↗

Diagnostic assessment of reservoir response to fracturing: a case study from Hydraulic Fracturing Test Site (HFTS) in Midland Basin

Abstract This paper outlines a data collection and diagnostics case study involving multiple horizontal shale wells. We look at well production profiles using rate transient analysis, differences in near wellbore complexity, geologic variations within the area of interest, as well as compositional differences in the rocks based on cores obtained from within the stimulated reservoir. The Hydraulic Fracturing Test Site is a multi-well experiment involving stimulation of unconventional shale wells in the southeastern Midland portion of the Permian Basin. The targeted formations include both the upper as well as the middle Wolfcamp formations, also referred alternatively as Wolfcamp A and Wolfcamp B. Data integration and analysis shared in this paper help us understand the various geologic controls impacting well productivity, particularly the wide variance observed between the Wolfcamp A and Wolfcamp B formations. Rate transient analysis indicates similar system permeabilities for stimulated wells. However, we observe higher effective fracture half-lengths for upper Wolfcamp wells. Using observations from 3D seismic interpretations (such as pad scale faults) as well as petrophysical and image log data, we highlight the substantial differences in stimulation as we move along the well laterals from the heel toward the toe sections. These differences are further reconciled with observations from zones with high data density at the core locations through stimulated rock, as well as independent data such as microseismic emissions. At the test site, Wolfcamp A was found to be relatively quartz rich with significant heterogeneity whereas Wolfcamp B is richer in clay and organic content. This impacts the geomechanical characteristics of the rock mass with much higher natural fracture density in the shallower interval. Thus, the fracture growth is more uniform in the deeper interval and more heterogeneous with branching likely in upper interval. Increased complexity also leads to consistently better productivity from the wells in the shallower interval as demonstrated from RTA results. This case study is unique because it provides valuable insights from actual sampling of the stimulated zones in hydraulically fractured wells and helps understand impact of various factors that contribute toward variability in well production. The findings from this study provides insights into need for optimization of completion designs in the various Wolfcamp landing zones, such as optimization of cluster or fracture spacing in various Wolfcamp intervals. In addition, it provides a useful template for data collection and research direction in future field test sites of similar nature in unconventional reservoirs.

Energy & Fuels↗

Open Energy Data Initiative (OpenEDI) - Open Data Access Tools (OEDI-ODAT) [SWR 20-57]

Open Energy Data Initiative (OpenEDI): Advancing Analytics and Research Innovation through Improved Data Access. The Open Energy Data Initiative (OEDI) aims to improve and automate access of high-value energy data sets across the U.S. Department of Energy’s (DOE’s) programs, offices, and national laboratories. Sponsored by the DOE, this platform is being implemented by the National Renewable Energy Laboratory (NREL) to make data actionable and discoverable by researchers and industry to accelerate analysis and advance innovation. Partners at other national laboratories will help utilize the platform for analysis, can be directly involved in determining requirements, and will be resources for providing new datasets that can be shared with the public to expand innovation.

Rager, David↗

Enriched Background Isotope Study (EBIS): Analysis of 14C-Enriched Carbon Cycle in Soils and Litter at Forested Oak Ridge and AmeriFlux Sites, 2001-2011

These data provide a record of the multi-year, multi-institutional Enriched Background Isotope Studies (EBIS) projects that ran from 2000 through 2011. Elevated levels of 14C enriched CO2 in the air and soil atmosphere as well as leaf, stem, and root tissues were observed on the Oak Ridge Reservations (ORR) during the summer of 1999, and were attributed to local incinerator activities on and/or near the ORR (Trumbore et al. 2002). The isolated enrichment of the background levels of 14C in local forest ecosystem represented a unique opportunity to study unresolved carbon cycling processes such as the contribution of leaf versus root litter contributions to soil carbon accumulation, the rate of vertical transport of carbon into deep soil storage pools, and the differential contribution of physicochemical versus faunal driven processes to soil carbon cycling and sequestration. Leaf litter from the local enriched forest was transplanted to selected sites on the ORR and to selected AmeriFlux study sites to study soil C cycling across a range of soils and climates. The EBIS research projects provide data on C flux from litter sources to mineral soil sinks for United States eastern hardwood forests necessary for testing process hypotheses and judging efficacy of soil C cycling models. Experimental results from this study are being used to parameterize and refine existing carbon dynamics models, the quantification of the long-term fate of ecosystem carbon inputs and as a means to judge the potential for ecosystem carbon sequestration via enhance litter inputs to soil. EBIS observations support conclusions that intra- and inter-annual soil carbon cycling in hardwood forest soils should be characterized as a least a two-compartment system where surface leaf-litter and belowground root turnover represent primary carbon sources for organic-layer and mineral-soil carbon cycles, respectively. EBIS experiments were conducted to complete enriched litterfall maniplations in upland forests on Ultisol and Inceptisol soils of the Oak Ridge Reservation, Oak Ridge, Tennessee. We also collected additional14C-enriched materials for new experimental applications, and applied those materials to multiple AmeriFlux sites over a range of climatic, edaphic and biological conditions. The research provided data for addressing DOE's goal of understanding mechanisms controlling C flux, and for the improvement of models to be applied to policy discussions regarding the safe levels of greenhouse gases for the earth's system. There are 5 data files provided in comma separated (*.csv) format for vegetation, field litter, soil and air [C] and C isotope data from the EBIS studies and associated environmental data.

54 ENVIRONMENTAL SCIENCES↗

FunM2C: A Filter for Uncertainty Visualization of Multivariate Data on Multi-Core Devices

Uncertainty visualization is an emerging research topic in data visualization because neglecting uncertainty in visualization can lead to inaccurate assessments. In this paper, we study the propagation of multivariate data uncertainty in visualization. Although there have been a few advancements in probabilistic uncertainty visualization of multivariate data, three critical challenges remain to be addressed. First, the state-of-the-art probabilistic uncertainty visualization framework is limited to bivariate data (two variables). Second, existing uncertainty visualization algorithms use computationally intensive techniques and lack support for cross-platform portability. Third, as a consequence of the computational expense, integration into production visualization tools is impractical. In this work, we address all three issues and make a threefold contribution. First, we take a step to generalize the state-of-the-art probabilistic framework for bivariate data to multivariate data with an arbitrary number of variables. Second, through utilization of VTK-m’s shared-memory parallelism and cross-platform compatibility features, we demonstrate acceleration of multivariate uncertainty visualization on different many-core architectures, including OpenMP and AMD GPUs. Third, we demonstrate the integration of our algorithms with the ParaView software. We demonstrate the utility of our algorithms through experiments on multivariate simulation data with three and four variables.

Hari, Gautam↗

Dataset: Breaking the barrier of human-annotated training data for machine-learning-aided plant research using aerial imagery

This dataset supports the implementation described in the manuscript "Breaking the Barrier of Human-Annotated Training Data for Machine-Learning-Aided Biological Research Using Aerial Imagery." It comprises UAV aerial imagery used to execute the code available at https://github.com/pixelvar79/GAN-Flowering-Detection-paper. For detailed information on dataset usage and instructions for implementing the code to reproduce the study, please refer to the GitHub repository.

generative and adversarial learning↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Dynamic Behavior of Natural Seep Vents: Analysis of Field and Laboratory Observations and Modeling (Final Scientific/Technical Report)

In this project, we have analyzed data collected by the U.S. Department of Energy (DOE), National Energy Technology Laboratory (NETL) in a high pressure water tunnel (HPWT) and data from two research cruises to natural seeps in the Gulf of Mexico to adapt and validate a numerical model to predict the dynamics of natural seeps in the deep oceans. The HPWT data include video observations of the shrinkage rate of individual methane and natural gas bubbles under simulated deep-water conditions. Field data were collected during two cruises by the Gulf Integrated Spill Research (GISR) Consortium led by Texas A&M University and funded by the Gulf of Mexico Research Initiative (GoMRI). These data included in situ observations from a remotely operated vehicle (ROV) of gas bubbles at two natural seep sites in the Gulf and acoustic observations of the natural seep bubble flares in the ocean water column. The acoustic data were from multibeam echosounders, one mounted in a forward-looking orientation on the ROV and another mounted down-looking in the haul of the ship. All of these laboratory and field data were focused on the dynamics of natural gas bubbles at temperatures and pressures favorable for clathrate hydrate formation between the gas and water. Our analyses of this data focused on understanding the mechanisms responsible for gas bubble dissolution within the hydrate stability zone (HSZ) of the oceans. Ice-like hydrate shells may form on the bubble-water interface under these conditions, and it was unknown how this might affect the mass transfer of gas into the ocean. We were able to extract bubble shrinkage rates from the HPWT datasets. Using this data we determined that mass transfer coefficients with and without a hydrate shell match empirical values for bubbles in contaminated systems (so-called dirty bubbles contaminated by naturally occurring surfactants). We also showed that free gas, and not gas hydrate, is the dominant dissolving phase when the hydrate sub-cooling is below 11 degree Celsius (temperature difference between hydrate the hydrate formation temperature and ambient temperator) or the pressure is reducing as bubbles rise through the ocean water column. Using this mass transfer model, our numerical model of bubble dissolution matched the over 200 HPWT experiments with an average error of 10% for predicting the bubble size at the end of an experiment. From field data in the literature, we also observed that gas bubbles dissolve faster when they are initially released, following mass transfer coefficients for so-called clean-bubbles (those not yet contaminated by surfactants). Shortly after release within the HSZ, a hydrate shell forms on the bubble-water interface, and the mass transfer reduces to rates matching those of dirty bubbles. We correlated this transition time from clean to dirty bubble behavior with the initial bubble surface area and the hydrate sub-cooling. With this model for hydrate formation time and using the mass transfer coefficients deduced from the HPWT data, we validated our numerical model for predicting the rise heights of natural seep flares in the oceans. Flare heights are commonly observed in haul-mounted acoustic multibeam data. The numerical model predicts bubbles to rise high in the ocean water column owing to the slower mass transfer rates for dirty bubbles that accompany the majority of their rise time. We found that the numerical model predictions matched the observed flare heights within 5% to 10% accuracy when we compared the rise heights of the largest bubbles released from the seafloor with the bubbles acoustically visible in the multibeam data. Bubbles become acoustically transparent as they shrink to sizes of order 1 mm in diameter for the multibeam frequencies used in the field. The forward-looking multibeam on the ROV also provided data on the lateral spreading of bubbles in natural seep flares. Our analysis of this data showed that spreading follows a diffusion process, with the effective diffusivity correlating with the wobbling length scale of these ellipsoidal bubbles. When we apply this diffusivity in a random displacement model of bubble spreading, our numerical simulations match closely the lateral spread observed by the M3 in the ocean water column. Finally, we compared the seep model predictions for the acoustic properties of these natural seep plumes with that observed by the acoustic instruments in the field. The M3 and EM 302 observations were converted to relative values of target strength using a calibration we obtained in the laboratory for the M3 and using an algorithm from the manufacturer for the EM 302. Comparing the numerical seep model to these data, we obtain good agreement over the whole height of rise of these bubble flares. This further validates the numerical model. Overall, our validated seep model captures the key dynamics of gas bubbles released from natural seeps in the oceans and helps to predict the fate of methane in the water column.

03 NATURAL GAS↗

Database of Nonaqueous Proton-Conducting Materials

This work presents the assembly of 48 papers, representing 74 different compounds and blends, into a machine-readable database of nonaqueous proton-conducting materials. SMILES was used to encode the chemical structures of the molecules, and we tabulated the reported proton conductivity, proton diffusion coefficient, and material composition for a total of 3152 data points. The data spans a broad range of temperatures ranging from -70 to 260 °C. To explore this landscape of nonaqueous proton conductors, DFT was used to calculate the proton affinity of 18 unique proton carriers. The results were then compared to the activation energy derived from fitting experimental data to the Arrhenius equation. It was found that while the widely recognized positive correlation between the activation energy and proton affinity may hold among closely related molecules, this correlation does not necessarily apply across a broader range of molecules. This work serves as an example of the potential analyses that can be conducted using literature data combined with emerging research tools in computation and data science to address specific materials design problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Developing Novel Performance Measures for Traffic Congestion Management and Operational Planning Based on Connected Vehicle Data

In this study, the authors present their efforts in exploring a new type of traffic data, referred to as internet-connected vehicle (ICV) data, for traffic congestion management and operational planning. Most currently manufactured vehicles contain onboard GPS and cellular modules, and they constantly connect to automobile manufacturers' clouds via cellular networks and upload their status. Some automobile manufacturers have recently redistributed the nonpersonal part of such data, such as geolocation, to third-party organizations for innovative applications. Compared with the traditional vehicle GPS data, the ICV data contain high-resolution GPS waypoints accompanied with the vehicles' abnormal moving events (e.g., hard braking). The ICV data also have huge potential in congestion management and operational planning. They explore to identify and analyze traffic congestion on both freeways and arterials using the ICV data. The ICV data adopted for this research are redistributed by Wejo Data Service, representing 10%-15% of all moving vehicles in the Dallas-Fort Worth (DFW) area in Texas. Through one case study for a freeway segment and one for an arterial segment, new traffic performance metrics based on the characteristics of ICV data have been presented. The highlights of these efforts are as follows: (I) queue length and propagation at freeway bottlenecks can be directly measured based on where and when most internet-connected vehicles slow down and join the queue; (II) an internet-connected vehicle's actual delay time on arterials can be directly measured according to its slow movement percentage, without assuming the nondelay travel speed; and (III) the ICV data set are also combined with the high-resolution traffic signal events to generate a ground-truth time-space diagram (TSD) on arterials - a common visualization of arterial signal performance for transportation planning and operations.

33 ADVANCED PROPULSION SYSTEMS↗