Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

DOE EV Data Collection - Facility Data

Facility data includes information on electricity consumption by larger-scale infrastructure, including buildings, solar arrays, and energy storage systems. Parameter definitions can be found in the data dictionary. If a connection between specific vehicle information and facility data exists, it will be available in the vehicle attributes table. Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data Format and Descriptions for the Alabama Carbon Storage: Data Sharing and Engagement Project

The Alabama Carbon Storage: Data Sharing and Engagement (ACS-DSE) project seeks to develop publicly accessible geologic carbon storage models and data across the southern Gulf Coastal Plain of Alabama. The public online platform developed for this project will include geologic, geophysical, infrastructure, and other relevant datasets and geologic models of the study area. Datasets, model surfaces (e.g. structural contour maps, isolith maps, porosity maps), and infrastructure data (e.g. offshore pipelines, field boundaries) will be downloadable in commonly used file formats. The anticipated primary geologic datasets are well headers, formation tops, average reservoir properties, and core analyses; these will be available as commaseparated values (CSV) text files and MS Excel workbooks. Geophysical logs will be available in Log ASCII Standard (LAS) file format. Modeled surfaces, such as structure contour maps, will be available in ArcGIS formats and text files. Infrastructure data will be available as ArcGIS shapefiles. This document provides information on the data sources and attributes of the datasets.

01 COAL, LIGNITE, AND PEAT

Data Centers and Digital Assurance Introduction to Supply Chain and Cybersecurity for Data Centers, Session 1

The first session of the TADA (Technical Assistance for Digital Assurance) Data Centers Cohort Workshop, held on October 30, 2025, introduced foundational concepts of Digital Assurance in the context of data center and grid integration. Sponsored by the U.S. Department of Energy, the workshop brought together utilities, data center operators, developers, and vendors to address cybersecurity and supply chain vulnerabilities. The session emphasized the growing criticality of data centers within the electric grid and the need for secure, real-time, bidirectional communication. Participants explored the principles of Digital Assurance, including cybersecurity, cyber-informed engineering (CIE), and lifecycle security, and applied a threat-vulnerability-consequence framework to identify and mitigate risks at the data center–grid interface. Discussions covered a range of threats such as spoofed dispatch signals and insider threats, architectural vulnerabilities like SCADA interfaces and insecure protocols, and potential consequences including cascading grid failures. The session also raised strategic questions about business value, vendor assurance, and defining cyber boundaries and responsibilities. This foundational workshop set the stage for deeper technical analysis and the development of actionable frameworks in subsequent sessions. Session 1 of 3.

24 - POWER TRANSMISSION AND DISTRIBUTION

Evaluation of Digital Nautical Chart data for confirmation and expansion of GeoNames data

Here, this work examines how Digital Nautical Chart (DNC) data may contribute to the evolution and refinement of GeoNames data for near-shore features. GeoNames features are point data with one or more possible place names. DNC Earth Cover Text (ECRText) objects are map labels positioned nearby their real word counterpart. ECRText feature map position strikes a compromise between association with real features and cartographic readability. This work explores whether ECRText features can confirm (or expand names for) existing locations or contribute new locations through data conflation. Due to name variations and spatial position, conflating these data are nontrivial. Previous work engaged in a brief examination using the trigram string matching algorithm under coarse proximity constraints, indicating that ECRText could provide additional value to GeoNames. This work builds on that study, by engaging in a deeper examination of spatial proximity and exploring conflation agreement across an ensemble of string matching approaches. The result finds strong ensemble agreement about ECRText features which already exist in GeoNames but mixed results about which features contribute new information, as well as exploring why some of these matching techniques fail. With an eye toward automation, computational efficiency was found not to be a constraint in sustaining updates.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression

Handling large-scale scientific data in high-performance computing (HPC) environments poses significant challenges, including excessive I/O, high storage costs, and slow query performance. Traditional approaches often require full data decompression and scans, making them impractical for real-time or interactive analysis. To address these limitations, we introduce Eureka, a unified data-index co-compression framework that enables fine-grained access and efficient range queries on compressed scientific datasets. Eureka integrates spatial domain decomposition with block-wise error-bounded lossy compression to support selective decompression. It constructs a hierarchical AVL-tree index during compression to capture block-level value ranges, enabling fast pruning during query execution. To reduce metadata overhead, the index itself is also compressed while ensuring recall-preserving results. Experiments on six diverse HPC simulation datasets show that Eureka achieves up to 25x data compression and over 300x index compression, surpassing state-of-the-art compressors such as SZ3 and ZFP in rate-distortion performance. Additionally, Eureka delivers over 30x speedup for low-selectivity range queries, making it a scalable and efficient solution for modern scientific data analysis.

Yan, Ning

Filling in Subsurface Storage Open Data Gaps - Updates to CCS Data Availability on EDX and EDX Spatial (FWP-1022465)

There is a need to preserve and efficiently access data resources to drive the next generation of research and development while ensuring compliance with DOE regulations. Over the last 10+ years, there has been ongoing efforts by the DOE Carbon Storage Program to ensure that there is effective data curation and preservation of DOE funded research leveraging the NETL-FECM data repository, the Energy Data eXchange (EDX). This talk presents updates about ongoing efforts to continue to support the mission of ensuring that carbon storage data is findable, accessible, interoperable, and reusable to the carbon storage stakeholder community through EDX and EDX Spatial. Presented at the NETL Carbon Management Review Meeting, Pittsburgh, 2024.

Morkner, Paige

Resolving the Discontinuity Between Thermal Scattering Data and Fast Data

The task of resolving discontinuities between thermal cross sections and fast cross sections in AMPX processing is part of a broader project to provide guidance on the potential effect of unknown thermal neutron scattering law (TNSL) data, primarily from the standpoint of criticality safety. The primary goal of this work is to deliver code fixes in AMPX that increase confidence that SCALE libraries provide physically correct cross sections across the full range of neutron energies encountered in transport applications. This report discusses the resolution of two sources of discontinuity. The first issue occurs when the Bragg edges in an evaluated nuclear data file do not extend to 5 eV. The second issue concerns parameters chosen during the processing of the nuclear data library—specifically, the parameter that defines the energy at which the thermal data ends and the fast nuclide data begins. This report also discusses a third issue, which is related to the calculation of probability distributions from a densely gridded double-differential cross section.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

47 Tuc in Rubin Data Preview 1. Exploring Early LSST Data and Science Potential

We present analyses of the early data from Rubin Observatory’s Data Preview 1 (DP1) for the field of the globular cluster 47 Tuc. The DP1 data set for 47 Tuc includes four nights of observations from the Rubin Commissioning Camera (LSSTComCam), covering multiple bands (ugriy). We address challenges of crowding in the inner region of the cluster and toward the SMC in DP1, and demonstrate improved star–galaxy separation by fitting fifth-degree polynomials to the stellar loci in color–color diagrams and applying multidimensional sigma clipping. We compile a catalog of 3576 probable 47 Tuc member stars selected via a combination of isochrone, Gaia proper-motion, and color–color space matched filtering. We explore the sources of photometric scatter in the 47 Tuc color–color sequence, evaluating contributions from various potential sources, including differential extinction within the cluster. Finally, of the 72 well-characterized variables in the field, we recover three known variable stars, including two RR Lyrae and one eclipsing binary, in the coadd-based object catalog, and identify 62 in the difference image-based object catalog. Although the DP1 lightcurves have sparse temporal sampling, they appear to follow the patterns of densely sampled literature lightcurves well. Despite some data limitations for crowded-field stellar analysis, DP1 demonstrates the promising scientific potential for future LSST data releases.

Choi, Yumi [NSF National Optical-Infrared Astronom

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark -- A Bayesian Inverse UQ-based Approach for Data Assimilation

The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. When comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that GLLS predictions fail to replicate computed response distributions for nonlinear applications, while MOCABA shows near agreement, and IUQ uses computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

FOS: Computer and information sciences

Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization

Domain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity—namely, the development of visualizations that support multiple groups. This development choice presents a trade-off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop Guidepost, a notebook-embedded visualization of data that helps scientists assess compute wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and design domains, applying them to categorize tasks based on their uniqueness across stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization through real-world case studies and a task-focused evaluation with nine participants. We observe that together, Guidepost's visual encodings, interactions, and export capabilities support the tasks of our differing personas.

97 MATHEMATICS AND COMPUTING

DOE EV Data Collection - Maintenance Data

Maintenance data includes information on maintenance performed on the electric vehicles, including preventive maintenance, service calls, and availability of the vehicles. The parameters collected, and their definitions, will vary due to the differences in maintenance tracking systems that exist between fleets. Parameter definitions are detailed in the data dictionary, and specific vehicle information is available in the vehicle attributes table. Vehicle ID can be used as a key between maintenance data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Advanced Data Center Energy Opportunities: Cloud and Infrastructure CoP - Data Center Energy and Efficiency with AI Adoption

The NLR portion of the "Cloud & Infrastructure CoP - Data Center Energy and Efficiency with AI Adoption" web meeting will cover data center locations, energy use and load growth, best practices, performance metrics, transition to direct liquid cooled data center equipment, and NLR's approach to optimizing data center.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data for Autonomous Transportation Awareness: Data Exchange Use Cases, Standards, and Barriers

This report examines the critical data exchanges between automated vehicle (AV) service providers and the cities and municipalities they serve. It assists municipal authorities in navigating the often complex and real-time digital data exchanges needed to support AV mobility services, with emphasis in three areas: (1) critical safety data for broad-area situational awareness of hazards typically associated emergency dispatch or roadway work zones; (2) performance metrics of AV services that inform the quantity, quality, spatial extents, and impact on the roadway network; and (3) regulatory and policy information, particularly dynamic information that governs how AV services interact with the roadway network, with emphasis on curb space. The report reviews existing practices and emerging protocols and standards and identifies key gaps to address moving forward.

33 ADVANCED PROPULSION SYSTEMS

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin

Circumventing data imbalance in magnetic ground state data for magnetic moment predictions

Abstract Magnetic materials play a crucial role in the transition to more sustainable forms of energy and electric vehicles. There is an anticipated shortage in magnetic materials in the future, and as a result there is an urgent need to discover and design new magnetic materials. Computational magnetic material design using density functional theory is daunting because of the challenge in identifying magnetic ground states from a combinatorially large set of possibilities. Machine learning offers a path forward by enabling efficient surrogate models that can more readily enumerate these states, but there is a dearth of training data available, and what is available tends to be imbalanced with too much non-magnetic data. In this work we show that the discrete and previously tackled data imbalance that exists at the level of the magnetic ordering leads to an imbalanced continuous distribution with many zeros when the data is unraveled at the atomic magnetic moment level, which subsequently leads to models with low accuracy for magnetic properties. We mitigate this by using a two-part model framework. Our scheme is able to classify atoms into magnetic and non-magnetic with an F1 score and Matthew’s correlation coefficient (MCC) of ~91% and then to provide an implicit embedding representation that maps directly onto the magnitude of the magnetic moment with a mean absolute error of 0.1 μ B . Beyond screening for new magnetic materials, we demonstrate an additional practical use case of our scheme: the provision of good initial guesses for magnetic moments in first-principles electronic relaxations. Such initialization is shown to lead to faster convergence to configurations that lie closer to the ground state.

Computer Science