Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

CORAL: A framework for rigorous self-validated data modeling and integrative, reproducible data analysis

Abstract Background Many organizations face challenges in managing and analyzing data, especially when relevant datasets arise from multiple sources and methods. Analyzing heterogeneous datasets and additional derived data requires rigorous tracking of their interrelationships and provenance. This task has long been a Grand Challenge of data science and has more recently been formalized in the FAIR principles: that all data objects be Findable, Accessible, Interoperable, and Reusable, both for machines and for people. Adherence to these principles is necessary for proper stewardship of information, for testing regulatory compliance, for measuring the efficiency of processes, and for facilitating reuse of data-analytical frameworks. Findings We present the Contextual Ontology-based Repository Analysis Library (CORAL), a platform that greatly facilitates adherence to all 4 of the FAIR principles, including the especially difficult challenge of making heterogeneous datasets Interoperable and Reusable across all parts of a large, long-lasting organization. To achieve this, CORAL's data model requires that data generators extensively document the context for all data, and our tools maintain that context throughout the entire analysis pipeline. CORAL also features a web interface for data generators to upload and explore data, as well as a Jupyter notebook interface for data analysts, both backed by a common API. Conclusions CORAL enables organizations to build FAIR data types on the fly as they are needed, avoiding the expense of bespoke data modeling. CORAL provides a uniquely powerful platform to enable integrative cross-dataset analyses, generating deeper insights than are possible using traditional analysis tools.

97 MATHEMATICS AND COMPUTING↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

Hydropower Capital and O&M Costs: An Exploration of the FERC Form 1 Data

This report explores the potential for using responses from the Federal Energy Regulatory Commission’s (FERC’s) “Electric Utility Annual Report,” also known as FERC Form 1, as a cost database for conventional and pumped storage hydropower (PSH). The report outlines the process used to compile an easy-to-access database from the original FERC Form 1 data and discusses the historical cost/performance data. The steps for developing the database include downloading the annual FoxPro files from the FERC website (FERC, 2020), reading them into an Excel format, and reconciling naming issues across years, repeated entries, and shared assets, among others. The cleaned and compiled Form 1 database is available on Oak Ridge National Laboratory’s (ORNL’s) HydroSource website. For this report, the compiled Form 1 database was linked to two other hydropower databases, the National Inventory of Dams and the ORNL Existing Hydropower Assets, that provide additional information on the characteristics and other features of plants in the Form 1 data. Explorations of the plant characteristics, performance, capital costs, and operation and maintenance (O&M) cost data in the combined database were presented separately for conventional hydropower and PSH projects. Form 1 has several advantages as a hydropower cost database. First, the cost data is reported directly by plant owners. Second, Form 1 provides cost breakdowns of capital and operating costs for large conventional and PSH plants enabling more detailed tracking of hydropower costs compared with typical total cost estimates. Third, although the Form 1 data is not reported by all operational hydropower plants in the United States, the available data represents a wide range of plant characteristics and a sizable proportion of the hydropower fleet (61% of PSH plants and 22% of conventional hydropower plants by capacity). The lower proportion of the conventional hydropower fleet is because Form 1 reporting requirements apply only to private utilities that meet given size thresholds, leaving out federally owned facilities and many smaller hydropower plants. Overall, the Form 1 cost database represents a unique, publicly available database on hydropower asset capital and O&M costs. This report leads to a compiled database for the reporting years 1994 to 2020 that helps resolve issues with accessibility and use of the FERC Form 1 data by hydropower plant owners, project developers, technology developers, and regulators.

13 HYDRO ENERGY↗

Enabling Floating Solar Photovoltaic (FPV) Deployment: FPV Technical Potential Assessment for Southeast Asia

Southeast Asia (SE Asia) is a region with growing energy demand and increasing development of floating solar photovoltaic (FPV) systems, which can help meet countries' renewable energy and energy security goals. This study uses a high-level geospatial assessment methodology to estimate the technical potential for monofacial and bifacial FPV on reservoirs and natural waterbodies in the ten countries within the Association of Southeast Asian Nations (ASEAN). Technical potential consists of the suitable waterbody area for FPV development (km2), the capacity of FPV that could be installed on this suitable area (MW), and the annual energy that could be generated from these installations (GWh/year). This first-of-its-kind FPV technical potential assessment for SE Asia can help policymakers and planners better understand the role that FPV could play in meeting regional energy demand and can ultimately guide investment decisions. Although this work focuses on SE Asia, the methodology may also be applicable for countries in other regions, with adaptations. The FPV technical potential results are also integrated into the Renewable Energy (RE) Data Explorer online tool (https://www.re-explorer.org/).

bifacial↗

Integrating Immersive Visualization in Molten-Salt Reactor Waste Management for Experimental Design and Planning

Molten-salt reactors (MSRs) represent a promising solution for next-generation nuclear energy, offering advantages in safety, fuel efficiency, and waste minimization. However, their liquid-fueled design presents unique challenges for spent fuel management, making post-shutdown waste characterization essential for developing effective strategies. Despite this need, there is a notable absence of visualization platforms specifically tailored to the unique characteristics and analytical requirements of MSR waste management. Existing tools in the nuclear industry are primarily designed for reactor operations or generic data exploration and lack both integration with MSR-specific multiphysics frameworks and the ability to simultaneously visualize time-dependent thermal fields, chemical composition evolution, and radiation distribution patterns. To address these limitations, this paper presents an immersive virtual reality (VR) visualization platform that processes and displays high-fidelity multiphysics simulation output from the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework in real-time, using Unity. The platform visualizes MSR waste characteristics such as nuclide decay, salt cooling, and corrosion by using Exodus II output data and running on a VR headset. It includes a user-friendly interface with features such as visibility toggling, cross-sectional slicing, and time-series animation for exploring simulation data. These capabilities support experimental design, stakeholder engagement, and public communication by making complex reactor behavior more accessible and understandable. By enhancing spatial reasoning and reducing cognitive load, this immersive environment fosters more effective communication and decision-making in MSR waste management.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data and Tools for Energy Planning and Analysis

The National Renewable Energy Laboratory (NREL) creates widely used data and tools to facilitate energy system planning and analysis. These software tools have been developed for complex research problems and perfected over real-world applications and laboratory validations. Some tools are award winners, others are open-source data explorers, and all are rigorously designed to empower decision-makers with accurate and accessible information. This software selection shows how NREL resources can help stakeholders achieve a clean, just, and resilient energy transformation.

data↗

Visualization for Scientific Discovery, Decision-Making, and Communication

Visualization–the use of visual elements to explore data, form hypotheses, or convey conclusions–is an integral part of the scientific process. Starting from an initial exploration of new data to illustrating outcomes to the general public, visualization is one of the most intuitive and powerful modes of communication. With the explosion of new data sources and types, unprecedented volumes of data, and new technologies, such as virtual reality and AI, visualization has become increasingly essential but also ever more challenging. Department of Energy’s (DOE) Office of Advanced Scientific Computing Research (ASCR) sponsored a Basic Research Needs workshop in January 2022 to understand the major opportunities and grand challenges in visualization tools and technologies for scientific computing, with a special focus on DOE-relevant applications and goals. The workshop identified five priority research directions (PRDs) for visualization to support scientific discovery, decision-making, and communication.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on Visualization for Scientific Discovery, Decision-Making, and Communication

Visualization—the use of visual elements to explore data, form hypotheses, or convey conclusions—is an integral part of the scientific process. Starting from an initial exploration of new data to illustrating outcomes for the general public, visualization is one of the most intuitive and powerful modes of communication. With the explosion of new data sources and types, unprecedented volumes of data, and new technologies, such as virtual reality (VR) and artificial intelligence (AI), visualization has become increasingly essential but also ever more challenging. The Department of Energy’s (DOE) Office of Advanced Scientific Computing Research (ASCR) sponsored a Basic Research Needs workshop in January 2022 to understand the major opportunities and grand challenges in visualization tools and technologies for scientific computing as well as for DOE-relevant applications and goals in general. The workshop identified five priority research directions (PRDs) for visualization to support scientific discovery, decision making, and communication. The first three PRDs describe interconnected research themes addressing the need for new techniques to deal with complex data, uncertainty, and interpretability (PRD 1); the need for scalable and interoperable software stacks (PRD 2); and the challenges and opportunities inherent in new technologies, such as VR, cloud, or exascale computing (PRD 3). The remaining two PRDs describe foundational research themes that recognize the potential of visualizations to provide equitable access to information and to strengthen the scientific discourse (PRD 4); and the need to consider human factors when designing visualizations (PRD 5). Collectively, these PRDs form the pillars for a coherent, long-term research and development strategy in Visualization for Scientific Discovery, Decision-Making, and Communication in the context of the Office of Science’s mission scope.

97 MATHEMATICS AND COMPUTING↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access↗

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives. This paper will outline the development, integration, output, and efficacy of the AskGDR LLM, including adherence to scientific rigor through improvements designed to increase the accuracy of generated answers, avoid speculation, and provide proper references for all resources used.

access↗

Data reduction considerations for the burning velocity of spherical constant volume flames of R32 (CH 2 F 2 ) with air

Here, the present work explores data reduction techniques for the measurement of the laminar burning velocities of R32(CH 2 F 2 )-air mixtures using a constant volume combustion device, in which the pressure-time history is the only measured parameter. To allow clear assessment of the accuracy of the data reduction methods, the pressure-time histories used for analysis are synthetically generated via a detailed numerical simulation employing full kinetics and with and without an optically-thin radiation model. Various data reduction models are employed, including a two-zone model and two multi-zone models, and these are compared with the results from the burning velocity obtained from the output of the numerical simulation. The data reduction schemes are shown to be accurate if the same radiation model is employed in the data reduction as was used in the flame simulation to generate the pressure trace used for post-processing. If the incorrect radiation model is employed, however, the errors can be quite large. The effects of stretch, radiation, and different data post-processing methodologies are explored and the errors quantified. Stretch is shown to be important for the early stages and the selected data range that is used for extrapolation has a significant effect on the extrapolated burning velocity. However, with an appropriate choice of data considered for extrapolation, the prediction of the unstretched burning velocity can be quite accurate.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Navigating Uncertainty: Challenges in Visualizing Ensemble Data and Surrogate Models for Decision Systems

Uncertainty visualization plays a critical role in transforming ensemble simulation data into actionable insights by effectively communicating various dimensions of uncertainty within a system. The emergence of artificial intelligence-driven surrogate models trained on multirun ensemble data offers a transformative opportunity to replace computationally intensive simulations with fast estimates, enabling users to explore data spaces with unprecedented depth and interactivity. However, integrating ensemble data and surrogate models into decision-making workflows and tools introduces novel challenges for uncertainty visualization. These include reconciling and clearly communicating the unique uncertainties associated with ensembles and their surrogate model estimates, and leveraging these approximations to inform actionable decisions. This work explores these challenges in the context of high-dimensional data visualization, bridging discrete datasets with their continuous representations and addressing the complexities of systems that support iterative navigation between input and output spaces. We evaluate the role of uncertainty visualization in fostering intuitive, actionable interactions and identify critical hurdles in advancing this frontier of computational simulation.

97 MATHEMATICS AND COMPUTING↗

Revisit NGC 5466 tidal stream with Gaia , SDSS/SEGUE, and LAMOST

ABSTRACT By mining the data from Gaia Early Data Release 3, Sloan Digital Sky Survey/Sloan Extension for Galactic Understanding and Exploration Data Release 16, and Large Sky Area Multi-Object Fiber Spectroscopic Telescope Data Release 8, 11 member stars of the NGC 5466 tidal stream are detected and 7 of them are newly identified. To reject contaminators, a variety of cuts are applied in sky position, colour–magnitude diagram, metallicity, proper motion, and radial velocity. We compare our data to a mock stream generated by modelling the cluster’s disruption under a smooth Galactic potential plus the Large Magellanic Cloud (LMC). The concordant trends in phase space between the model and observations imply that the stream might have been perturbed by the LMC. The two most distant stars among the 11 detected members trace the stream’s length to 60° of sky, supporting and extending the previous length of 45°. Given that NGC 5466 is so distant and potentially has a longer tail than previously thought, we expect that the NGC 5466 tidal stream could be a useful tool in constraining the Milky Way gravitational field.

79 ASTRONOMY AND ASTROPHYSICS↗

High-Fidelity Solar Irradiance Data: Simple Access to State-of-the-Art Information Accelerates Southeast Asia's Clean Energy Economic Transformation

High quality, robust, and reliable renewable energy resource data is foundational to climate-smart decision making, evidence-based policy planning, and clean energy investment mobilization. USAID and NREL, through the Advanced Energy Partnership for Asia, are expanding access to this critical resource data by providing free, high-fidelity solar resource data for Southeast Asia through the RE Data Explorer platform. This brief highlights several ways the Southeast Asia solar resource data has been used for power system planning and project development in the region.

Advanced Energy Partnership for Asia↗

High-Resolution Southeast Asia Wind Resource Data Set

Well informed decision-making is a key part of integrating variable renewable energy into the global energy marketplace. USAID and NREL, through the Advanced Energy Partnership for Asia, are expanding access to critical resource data by providing free, high-fidelity time-series wind resource data for Southeast Asia through the RE Data Explorer platform. This brief highlights the development of the Southeast Asia wind resource data set and discusses the impacts of this data.

Advanced Energy Partnership for Asia↗