Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

First Plant Cell Atlas symposium report

The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.

59 BASIC BIOLOGICAL SCIENCES↗

The IsoGenie database: an interdisciplinary data management solution for ecosystems biology and environmental research

Modern microbial and ecosystem sciences require diverse interdisciplinary teams that are often challenged in “speaking” to one another due to different languages and data product types. Here we introduce the IsoGenie Database, a de novo developed data management and exploration platform, as a solution to this challenge of accurately representing and integrating heterogenous environmental and microbial data across ecosystem scales. The IsoGenieDB is a public and private data infrastructure designed to store and query data generated by the IsoGenie Project, a ~10 year DOE-funded project focused on discovering ecosystem climate feedbacks in a thawing permafrost landscape. The IsoGenieDB provides (i) a platform for IsoGenie Project members to explore the project’s interdisciplinary datasets across scales through the inherent relationships among data entities, (ii) a framework to consolidate and harmonize the datasets needed by the team’s modelers, and (iii) a public venue that leverages the same spatially explicit, disciplinarily integrated data structure to share published datasets. The IsoGenieDB is also being expanded to cover the NASA-funded Archaea to Atmosphere (A2A) project, which scales the findings of IsoGenie to a broader suite of Arctic peatlands, via the umbrella A2A Database (A2A-DB). The IsoGenieDB’s expandability and flexible architecture allow it to serve as an example ecosystems database.

54 ENVIRONMENTAL SCIENCES↗

Ascribe XR v0.1.0

Ascribe XR is an immersive visualization software designed for scientists and engineers working with 3D data sets. Its key features include interactive exploration, multi-user collaboration, and flexible data import capabilities, supporting various formats such as meshes, volumes, and terrain maps. The software utilizes Godot, OpenXR and PC-VR technology to provide an immersive experience. Ascribe XR is used for data analysis, visualization, and collaboration in various fields, enabling users to gain deeper insights into complex data sets. Its advantages over similar technologies include its flexibility, customizability, and ease of use. Ascribe XR's interactive and immersive environment facilitates collaboration and accelerates the discovery process. Compared to traditional 2D visualization tools, Ascribe XR offers a more engaging and intuitive experience, allowing users to explore complex data sets in a more natural and interactive way. Its ability to support multi-user collaboration and flexible data import capabilities make it a versatile tool for various applications. Overall, Ascribe XR provides a unique combination of features, usability, and performance, making it an attractive solution for scientists and engineers working with 3D data sets.

Pandolfi, Ronald [Lawrence Berkeley National Labor↗

Selection and Ranking of Experiments from the Halden Database in support of Multiscale Model Validation

Validating fuel performance codes, such as BISON, requires an extensive amount of experiments covering a wide range of operating conditions and fuel types. Within the light-water reactor (LWR) space, there have been several international experimental programs that have contributed to the wealth of available experimental data available for use. One of those international programs, the Halden Reactor Project (HRP), began in 1958 and utilized the Halden Boiling Water Reactor to conduct many highly instrumented experiments until the reactor closed in 2018. Idaho National Laboratory, through the U.S. Department of Energy, has utilized several Halden experiments to perform the initial validation of the BISON code based upon their inclusion in international modeling and simulation benchmarks. Recently, the HRP has provided member organizations a complete copy of all available data, reports, and presentations since the HRP began. This report provides an initial exploration of the data available in the database for use in validating the multiscale models under development in the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program for LWR applications. Ranking tables that identify potential validation cases are provided for the high priority models of interest. It was found that some of the recommended high priority experiments correspond to additional rods in existing assemblies already available in the BISON validation suite. It is expected that several of these cases will be incorporated into future NEAMS milestones in the fuels technical area for increased validation of BISON.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Adapting Grid Criticality for Data Centers

This presentation explores the evolving definition of “critical load” in the electric grid, emphasizing the growing importance of digital infrastructure—particularly data centers—in grid resilience, restoration, and modernization. As utilities increasingly rely on AI-driven analytics and software-defined control systems, data centers have shifted from passive electricity consumers to essential computational hubs that enable National Critical Functions (NCFs) and support real-time grid operations. The deck examines the scale and impact of digital loads, the need for grid modernization to manage rapid load growth, and the diverse computing paradigms required for AI deployment. It introduces a tiered taxonomy for classifying critical loads, highlights operational dependencies between the grid and digital infrastructure, and discusses policy implications for integrating data centers into emergency planning and restoration protocols. Through case studies and practical frameworks, the presentation provides actionable insights for utilities, regulators, and planners navigating the digital transformation of the power sector.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

Efficient data acquisition and training of collisional-radiative model artificial neural network surrogates through adaptive parameter space sampling

Abstract Effective plasma transport modeling of magnetically confined fusion devices relies on having an accurate understanding of the ion composition and radiative power losses of the plasma. Generally, these quantities can be obtained from solutions of a collisional-radiative (CR) model at each time step within a plasma transport simulation. However, even compact, approximate CR models can be computationally onerous to evaluate, and in-situ evaluation of these models within a larger plasma transport code can lead to a rigid bottleneck. As a way to bypass this bottleneck, we propose deploying artificial neural network (ANN) surrogates to allow rapid evaluation of the necessary plasma quantities. However, one issue with training an accurate ANN surrogate is the reliance on a sufficiently large and representative training and validation data set, which can be time-consuming to generate. In this work we explore a data-driven active learning and training routine to allow autonomous adaptive sampling of the problem parameter space to ensure a sufficiently large and meaningful set of training data is assembled for the network training. As a result, we can demonstrate approximately order-of-magnitude savings in required training data samples to produce an accurate surrogate.

97 MATHEMATICS AND COMPUTING↗

Exploiting Commonly-Reported Age-Hardening Data and Discovering Systematics Across Metallic Alloys and Alloy Systems to Identify Corrosion’s Most Influential Factors

This work explored ways to data mine legacy literature and predict solid-solid precipitation in support of simpler assessment of corrosion propensity in metallic alloys. Of interest was locating the peak age watershed (maximum hardness or strength), beyond which lies the regime termed “overaging.” A diligent search of literature and reference books, discussions with SMEs, and application of various algorithms showed that none of the premises going in held up to scrutiny.

36 MATERIALS SCIENCE↗

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology↗

Exploring physics of ferroelectric domain walls via Bayesian analysis of atomically resolved STEM data

The physics of ferroelectric domain walls is explored using the Bayesian inference analysis of atomically resolved STEM data. We demonstrate that domain wall profile shapes are ultimately sensitive to the nature of the order parameter in the material, including the functional form of Ginzburg-Landau-Devonshire expansion, and numerical value of the corresponding parameters. The preexisting materials knowledge naturally folds in the Bayesian framework in the form of prior distributions, with the different order parameters forming competing (or hierarchical) models. Here, we explore the physics of the ferroelectric domain walls in BiFeO 3 using this method, and derive the posterior estimates of relevant parameters. More generally, this inference approach both allows learning materials physics from experimental data with associated uncertainty quantification, and establishing guidelines for instrumental development answering questions on what resolution and information limits are necessary for reliable observation of specific physical mechanisms of interest.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Information-Theoretic Exploration of Multivariate Time-Varying Image Databases

Modern scientific simulations produce very large datasets, making interactive exploration of such data computationally prohibitive. An increasingly common data reduction technique is to store visualizations and other data extracts in a database. The Cinema project is one such approach, storing visualizations in an image database for post hoc exploration and interactive image-based analysis. This work focuses on developing efficient algorithms that can quantify various types of multivariate dependencies existing within multi-variable datasets. It applies specific mutual information measures for the quantification of salient regions from multivariate image data. Here, using such information measures, the opacity of the images is modulated so that the salient regions are automatically highlighted and the domain scientists can interactively explore the most relevant regions for scientific discovery.

97 MATHEMATICS AND COMPUTING↗

BASIN-3D: A brokering framework to integrate diverse environmental data

Diverse observational and simulation datasets are needed to understand and predict complex ecosystem behavior over seasonal to decadal and century time-scales. Integration of these datasets poses a major barrier towards advancing environmental science, particularly due to differences in the structure and formats of data provided by various sources. Here, we describe BASIN-3D (Broker for Assimilation, Synthesis and Integration of eNvironmental Diverse, Distributed Datasets), a data integration framework designed to dynamically retrieve and transform heterogeneous data from different sources into a common format to provide an integrated view. BASIN-3D enables users to adopt a standardized approach for data retrieval and avoid customizations for the data type or source. We demonstrate the value of BASIN-3D with two use cases that require integration of data from regional to watershed spatial scales. The first application uses the BASIN-3D Python library to integrate time-series hydrological and meteorological data to provide standardized inputs to analytical and machine learning codes in order to predict the impacts of hydrological disturbances on large river corridors of the United States. The second application uses the BASIN-3D Django framework to integrate diverse time-series data in a mountainous watershed in East River, Colorado, United States to enable scientific researchers to explore and download data through an interactive web portal. Thus, BASIN-3D can be used to support data integration for both web-based tools, as well as data analytics using Python scripting and extensions like Jupyter notebooks. The framework is expected to be transferable to and useful for many other field and modeling studies.

Varadharajan, C↗

Tiling Framework for Heterogeneous Computing of Matrix based Tiled Algorithms

Tiling matrix operations can improve the load balancing and performance of applications on heterogeneous computing resources. Writing a tile-based algorithm for each operation with a traditional, hand-tuned tiling approach that uses for loops in C/C++ is cumbersome and error prone. Moreover, it must enable and support the heterogeneous memory management of data objects and also explore architecture-supported, native, tiled-data transfer APIs instead of copying the tiled data to continuous memory before the data transfer. The tiling framework provides a tiled data structure for heterogeneous memory mapping and parameterization to a heterogeneous task specification API. We have integrated our tiled framework into MatRIS (Math kernels library using IRIS). IRIS is a heterogeneous run-time framework with a heterogeneous programming model, memory model, and task execution model. Experiments reveal that the tiled framework for BLAS operations has improved the programmability of tiled BLAS and improved performance by ~20% when compared against the traditional method that copies the data to continuous memory locations for heterogeneous computing.

Miniskar, Narasinga Rao↗

Peregrine Software Development: Report on the Code Conversion From Python to C++

This work package seeks to convert the Peregrine software tool from its original Python implementation to a production version based on the C++ language. Peregrine is a powerful research platform with a multitude of advanced data analytics and data visualization functionalities. Developed by scientists to explore multimodal and multidimensional data related to the production of components using powder bed additive manufacturing processes, the tool implements state-of-the-art algorithms to assist machine users in making build or part quality determinations. Given that Peregrine is data-intensive, the goal of this conversion is to enhance the tool’s flexibility and interactivity and reduce the number of code dependencies to facilitate its deployment as part of the ongoing technology transfer campaign. This brief document provides an overview of Peregrine’s functionalities and capabilities, along with a detailed description of the core functionalities that have been implemented to date in the new C++ version. This document serves as a development update at the end of the first year of the ongoing conversion and will be regularly updated as progress continues.

97 MATHEMATICS AND COMPUTING↗

Ensemble Kalman inversion of induced polarization data

SUMMARY This paper explores the applicability of ensemble Kalman inversion (EKI) with level-set parametrization for solving geophysical inverse problems. In particular, we focus on its extension to induced polarization (IP) data with uncertainty quantification. IP data may provide rich information on characteristics of geological materials due to its sensitivity to characteristics of the pore–grain interface. In many IP studies, different geological units are juxtaposed and the goal is to delineate these units and obtain estimates of unit properties with uncertainty bounds. Conventional inversion of IP data does not resolve well sharp interfaces and tends to reduce and smooth resistivity variations, while not readily providing uncertainty estimates. Recently, it has been shown for DC resistivity that EKI is an efficient solver for inverse problems which provides uncertainty quantification, and its combination with level set parametrization can delineate arbitrary interfaces well. In this contribution, we demonstrate the extension of EKI to IP data using a sequential approach, where the mean field obtained from DC resistivity inversion is used as input for a separate phase angle inversion. We illustrate our workflow using a series of synthetic and field examples. Variations with uncertainty bounds in both DC resistivity and phase angles are recovered by EKI, which provides useful information for hydrogeological site characterization. Although phase angles are less well-resolved than DC resistivity, partly due to their smaller range and higher percentage data errors, it complements DC resistivity for site characterization. Overall, EKI with level set parametrization provides a practical approach forward for efficient hydrogeophysical imaging under uncertainty.

Geochemistry & Geophysics↗

Capturing the Physics of MaNGA Galaxies with Self-supervised Machine Learning

As available data sets grow in size and complexity, advanced visualization tools enabling their exploration and analysis become more important. In modern astronomy, integral field spectroscopic galaxy surveys are a clear example of increasing high dimensionality and complex data sets, which challenges the traditional methods used to extract the physical information they contain. Here, we present the use of a novel self-supervised machine-learning method to visualize the multidimensional information on stellar population and kinematics in the MaNGA survey in a 2D plane. Our framework is insensitive to nonphysical properties such as the size of the integral field unit and is therefore able to order galaxies according to their resolved physical properties. Using the extracted representations, we study how galaxies distribute based on their resolved and global physical properties. We show that even when exclusively using information about the internal structure, galaxies naturally cluster into two well-known categories, rotating main-sequence disks and massive slow rotators, from a purely data-driven perspective, hence confirming distinct assembly channels. Low-mass rotation-dominated quenched galaxies appear as a third cluster only if information about the integrated physical properties is preserved, suggesting a mixture of assembly processes for these galaxies without any particular signature in their internal kinematics that distinguishes them from the two main groups. The framework for data exploration is publicly released with this publication, ready to be used with the MaNGA or other integral field data sets.

79 ASTRONOMY AND ASTROPHYSICS↗

Next Generation System Analysis Model Recently Added Features and Future Plans - Abstract

The Nuclear Waste Policy Act of 1982, as amended (NWPA 1982), established the federal government’s responsibility to accept spent nuclear fuel (SNF) and high-level radioactive waste (HLW) from waste owners and generators for ultimate disposition. SNF generated by the current fleet of commercial nuclear reactors is being stored at the reactor sites in spent fuel pools (SFPs) and in dry independent spent fuel storage installations (ISFSIs). The US Department of Energy Office of Nuclear Energy (DOE-NE) is developing an Integrated Waste Management Program (IWMP) comprising a suite of options and supporting analyses to enable future informed choices. The IWMP is applying integrated waste management system architecture analysis, system engineering, and decision analysis principles to inform potential future decisions regarding potential nuclear waste management system architectures. Architecture analyses of the IWM system are being conducted to support the future deployment of a comprehensive system for managing nuclear waste that considers all major aspects of the back end of the nuclear fuel cycle (i.e., transportation, storage, and disposal). The Next Generation System Analysis Model (NGSAM) is an agent-based simulation software tool designed for the express purpose of modeling the IWM system. NGSAM imports data from the Oak Ridge National Laboratory (ORNL) Unified Database (e.g., historic assembly information, thermal profiles for assembly heat, at-reactor dry storage loadings) to ensure that the simulation initializes with a realistic representation of the state of commercial SNF in the United States. Recent major enhancements that have been implemented into NGSAM since NGSAM was last presented at the WM2019 conference include: • Tracking of railroad escort and buffer car acquisition. • Addition of heavy haul and barge routes for some sites, as well as support for user-defined inter-modal routes. • Updates to the logic that checks the thermal maps prior to package transport. • Addition of an allocation method that predicts when reactor sites will pack assemblies from their pools for dry storage and allocates packages to those reactor sites in the preceding periods, favoring direct transport packages and reducing the number of packages that reactor sites pack for dry storage at their ISFSIs. • Addition of reactor site family operational limits, which are used to limit the number of loads from the pool and from dry storage at a given reactor site per year. • Support has been added for multiple canister loading maps and packages having multiple compatible transportation overpacks. • Updates in the handling of non-commercial fuel, including a new database containing data to support the updates. • Support for repackaging at reactor sites. • Implementing additional output reports or modifying existing reports. • User edits can now be created and edited via the NGSAM website. • Ability to load packages for dry storage at ISF pools. • Same-type package blending at DOE sites. • Support for multi-mode transloading at reactor sites. These new features have improved NGSAM capabilities and/or improve the user experience with the model and will be discussed in more detail. The initial NGSAM requirements for advanced reactor fuels, reprocessing, treatment, and conditioning are preliminary and are described at a high level in this paper: analysts will provide more specific requirements to the NGSAM team in the future. Additionally, there are many data needs associated with modeling advanced reactors in NGSAM, but many of the data or plans are still in progress and/or yet to be fully defined. However, this document describes an initial exploration of the data relevant to this program. Advanced reactor data will likely require revision as concepts evolve and new considerations are made. This is a technical paper that does not take into account contractual limitations or obligations under the Standard Contract for Disposal of Spent Nuclear Fuel and/or High-Level Radioactive Waste (Standard Contract) (10 CFR Part 961). For example, under the provisions of the Standard Contract, spent nuclear fuel in multi-assembly canisters is not an acceptable waste form, absent a mutually agreed to contract amendment. To the extent discussions or recommendations in this paper conflict with the provisions of the Standard Contract, the Standard Contract governs the obligations of the parties, and this paper in no manner supersedes, overrides, or amends the Standard Contract. This paper reflects technical work which could support future decision making by DOE. No inferences should be drawn from this paper regarding future actions by DOE, which are limited both by the terms of the Standard Contract and Congressional appropriations for the Department to fulfill its obligations under the Nuclear Waste Policy Act including licensing and construction of a spent nuclear fuel repository.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Unified Software Architecture for Advanced Materials and Manufacturing Technologies Data Management and Processing: FY 2023 Multidimensional Data Correlation Platform

This report details the various digital manufacturing activities ongoing at ORNL as part of the Advanced Materials and Manufacturing Technologies (AMMT) program. The AMMT program is exploring a data-driven approach to demonstrate the use of AM for the fabrication of components for nuclear applications, with the goal of providing a greater understanding of manufacturing quality outcomes that would pave the way toward the development of standards for certification and qualification. The objective of this work package is to establish a digital manufacturing discipline common to all participants of the AMMT program to improve the performance, reliability, and lifetime of nuclear components. As part of this effort, we will develop a unified software architecture for AMMT data management and processing, deploy the digital platform across AMMT participants’ facilities, and generate pedigreed datasets in a common format across multiple labs and facilities. To this end, the MDDC work package has focused on three activities during FY23. First, the MDF Digital Tool was overhauled to better serve the needs of the AMMT program. Next, multiple laser powder bed fusion (L-PBF) systems at the MDF were upgraded to a common sensor package for collecting comparable in situ data across machines. Finally, various improvements relevant to the AMMT program were implemented in the ORNL-developed software tool, Peregrine. This report marks the completion of FY23 milestone M3CR-22OR0403051: Report Describing the Architecture of the Digital Platform to Support AMMT Activities.

36 MATERIALS SCIENCE↗