Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

dCache project status and update

The dCache project delivers an open-source, massively scalable, distributed storage system deployed internationally to satisfy today’s scientists’ ever-demanding storage requirements. Its multifaceted approach supports different use cases with the same storage, from high throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and longterm data persistence on tertiary storage. Even though dCache was initially developed for HEP experiments, today, it is used by various scientific communities, including astrophysics, biomed, and life science, each with their specific requirements. To match the needs of these new communities and keep up with the scaling demands of existing experiments, dCache is permanently evolving. With this contribution, we would like to highlight the recent developments in dCache regarding integration with CERN Tape Archive (CTA), advanced metadata handling, token-based authorization support, bulk API for QoS transitions, REST API to control interaction with the tape system, and future development directions.

Mkrtchyan, Tigran [DESY]↗

Distribution Grid Model Publication Investigation

Interest in the external exchange of distribution grid model data is growing around the world, driven largely by the challenges and opportunities presented by the increasing amount of generation, storage, and flexible load being embedded within the distribution grid. This report provides an overview of the current state of distribution grid model data sharing, with a focus on the industry-leading activities currently underway in Great Britain (GB). A second report will explore opportunities for external distribution grid model sharing in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Conquering Data Chaos: Research Data Management with Kubernetes

Managing massive volumes of data and effectively making it accessible to researchers poses significant challenges and is a barrier to scientific discovery. In many cases, critical data is locked up in unwieldy file formats or one-off databases and is too large to effectively process on a single machine. This talk explores the role of Kubernetes, an open-source container orchestration platform, in addressing research data management challenges. I will discuss how we are using a set of publicly available open-source and home-grown tools in the National Renewable Energy Lab (NREL) Data, Analysis, and Visualization (DAV) group to help researchers overcome data-related bottlenecks. The talk will begin by providing an overview of the data challenges faced in research data management, including data storage, processing, and analysis. I will highlight Kubernetes' ability to handle large-scale data by leveraging containerization and distributed computing, including distributed storage. Kubernetes allows researchers to encapsulate data processing infrastructure and workflows into portable containers, enabling reproducibility and ease of deployment. Kubernetes can then schedule and manage the resource allocation of these containers to enable efficient utilization of limited computing resources, leading to more efficient data processing and analysis. I will discuss some limitations of traditional, siloed approaches to dealing with data and emphasize the need for solutions which foster collaboration. I will highlight how we are using Kubernetes at NREL to facilitate data sharing and cooperation among research teams. Kubernetes' flexible architecture enables the deployment of shared computing environments, such as Apache Superset, where researchers can seamlessly access and analyze shared datasets. Providing the ability to have one research team easily consume data generated by another, utilizing Kubernetes' as a central data platform, is one of the major wins we've encountered by adopting the platform. Finally, I will showcase real-world use cases from NREL where we have used Kubernetes to solve some persistent data challenges involving large volumes of sensor and monitoring data. I will discuss the challenges we encountered when creating our cluster and making it available as a production-ready resource. I will also discuss the specific suite of tools, including Postgres and Apache Druid for columnar and timeseries data, and Redpanda Kafka for streaming data we have deployed in our infrastructure, and the process that went into the selection of these tools.

collaborative environment↗

Data Request for the Distribution Grid Atlas

Model-based, distribution powerflow analysis is a foundational component of system planning and grid modernization efforts, but data security is an impediment to collaboration among utility engineers, researchers, developers, community members, and other stakeholders. Pacific Northwest National Laboratory (PNNL) and the National Renewable Energy Laboratory (NREL) are partnering to develop the new Distribution Grid Atlas - a publicly available catalog of realistic, geographically relevant, representative distribution feeder models without sensitive geographic information, customer data, or disclosure of utility models. We are looking for utilities to share data for the Distribution Grid Atlas.

distribution↗

dCache: Inter-disciplinary storage system

The dCache project provides open-source software deployed internationally to satisfy ever more demanding storage requirements. Its multifaceted approach provides an integrated way of supporting different use-cases with the same storage, from high throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters and long term data persistence on a tertiary storage. Though it was originally developed for the HEP experiments, today it is used by various scientific communities, including astrophysics, biomed, life science, which have their specific requirements. In this paper we describe some of the new requirements as well as demonstrate how dCache developers are addressing them.

Mkrtchyan, Tigran↗

DuraMAT Data Hub

The DuraMAT Data Hub has been supporting the consortium for the past six years. The Data Hub has had success in supporting the projects, providing a platform for sharing data within projects and to the public, and learning how to better leverage the existing software platform and the available Amazon Web Services environment. During this new generation of the Data Hub, we are looking at ways to help improve the data hub architecture, user experience, and improve operations by taking advantage of new technology platforms and software that will be more impactful on the consortium researchers and the broader scientific community. In this poster we will look at the current operational capabilities, data dissemination, and development that will improve the system in the near and far future.

14 SOLAR ENERGY↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Securing Environmental IoT Data Using Masked Authentication Messaging Protocol in a DAG-Based Blockchain: IOTA Tangle

The demand for the digital monitoring of environmental ecosystems is high and growing rapidly as a means of protecting the public and managing the environment. However, before data, algorithms, and models can be mobilized at scale, there are considerable concerns associated with privacy and security that can negatively affect the adoption of technology within this domain. In this paper, we propose the advancement of electronic environmental monitoring through the capability provided by the blockchain. The blockchain’s use of a distributed ledger as its underlying infrastructure is an attractive approach to counter these privacy and security issues, although its performance and ability to manage sensor data must be assessed. We focus on a new distributed ledger technology for the IoT, called IOTA, that is based on a directed acyclic graph. IOTA overcomes the current limitations of the blockchain and offers a data communication protocol called masked authenticated messaging for secure data sharing among Internet of Things (IoT) devices. We show how the application layer employing the data communication protocol, MAM, can support the secure transmission, storage, and retrieval of encrypted environmental sensor data by using an immutable distributed ledger such as that shown in IOTA. Finally, we evaluate, compare, and analyze the performance of the MAM protocol against a non-protocol approach.

Gangwani, Pranav (ORCID:0000000159226002)↗

Datashare

Datashare facilitates communication and data sharing within local networks in potentially dangerous situations such as an explosive ordnance disposal. During such events, there is a need to transmit information rapidly around the incident area. It is a distributed database that does not require an internet connection for operation. In addition, Datashare interfaces with XTK and other software applications, allowing for seamless integration and data management. Datashare supports video calls over the network, enabling real-time communication among users. This software serves to organize, package, and share between responders on location and export data to those off location. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Eldridge, Bryce [Sandia National Lab. (SNL-CA), Li↗

pnnl/UUDEX

UUDEX describes a communications architecture and protocol suite that allows organizations to exchange data and information. It does this by defining relations between a set of client nodes that share data via a set of server nodes. The primary focus of UUDEX is to facilitate communications between control centers, operations centers, and other trusted organization.

Welsh, Jeff↗

“Translational Opportunities in CPS Transportation”

This talk will describe opportunities for translating research from open-road experiments with modified adaptive cruise controllers. The data and controllers used for the previous research are based on use cases and example drives from within the US. We explore processes, data sharing, experiments, and other techniques that we think will drive trans-Pacific partnerships with driving data.

Sprinkle, Jonathan↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Best practices: Organizational execution

This article, the fourth and final installment in the DOE–IDEA series, focuses on organizational execution and how strong management practices, teamwork, and preparedness contribute to successful district energy systems. It highlights case studies from Ashley Energy and Cornell University to illustrate effective operational strategies. Ashley Energy demonstrates the importance of emergency preparedness and rapid response. After a major flood disrupted its plant, the organization restored service in under 72 hours by relying on pre-established plans, vendor relationships, and trained staff. The case emphasizes proactive contingency planning, understanding insurance processes, and empowering skilled personnel to improvise during crises. Cornell University’s example highlights the role of collaboration and transparency in long-term success. Its district energy system benefits from strong data sharing, real-time energy monitoring, and active involvement of faculty and students in system planning and innovation. This culture of teamwork and data-driven decision-making supports sustainability goals and continuous system improvement. Overall, the article shows that effective organizational execution—through preparedness, collaboration, and data transparency—is essential for maintaining reliable, efficient, and sustainable district energy systems.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Floating Offshore Wind US Manufacturing and Commercialization: Cooperative Research and Development (Final Report)

NREL assessed the supply chain and workforce considerations for the OCG-Wind floater technology, a floating semi-submersible offshore wind substructure, as well sharing vessel needs to inform their installation strategy. This technical assistance was in support of the FLoating Offshore Wind ReadINess (FLOWIN) Prize Phase 2 submission. NREL provided an assessment of domestic supplier capabilities for the main components of their floating offshore wind platform design and analyzed US regional and national supply chain constraints and gaps. Thirteen interviews with companies including steel distributors, forges, foundries, ports, large component fabricators, subcomponent fabricators, and secondary suppliers provided key insights such as 1) assembly ports are the key infrastructure barrier standing in the way of unlocking the domestic assembly and component fabrication for steel-based FOW platforms, 2) domestic steel producers can supply the types and quantities of steel necessary for FOW platforms, and 3) coordination between stakeholders will be a vital part of successfully developing the supply chain and infrastructure needed to domestically produce FOW platforms. In the workforce assessment, NREL documented a step-by-step approach to conduct a place-based assessment of the foundational workforce consideration for recruiting, upskilling, and retaining a workforce, such as supportive local and state policy, nearby education and training programs, and existing relevant industry. This approach was applied to Tacoma, Washington. Tacoma was indicated to have the potential be a successful location for fabrication and assembly of floating offshore wind energy in terms of workforce development. To share data on vessel requirements to install the OCG-Wind floater, NREL compiled resources that help answer the questions related to anchor handling tug vessels, shared a database of cable laying vessels, and answered questions on complying with the Jones Act.

17 WIND ENERGY↗

Earth and Space Science Informatics Perspectives on Integrated, Coordinated, Open, Networked (ICON) Science

Abstract This article is composed of three independent commentaries about the state of Integrated, Coordinated, Open, Networked (ICON) principles (Goldman, et al., 2021b, https://doi.org/10.1029/2021EO153180 ) in Earth and Space Science Informatics (ESSI) and includes discussion on the opportunities and challenges of adopting them. Each commentary focuses on a different topic: (Section 2) Global collaboration, cyberinfrastructure, and data sharing; (Section 3) Machine learning for multiscale modeling; (Section 4) Aerial and satellite remote sensing for advancing Earth system model development by integrating field and ancillary data. ESSI addresses data management practices, computation and analysis, and hardware and software infrastructure. Our role in ICON science therefore involves collaborative work to assess, design, implement, and promote practices and tools that enable effective data management, discovery, integration, and reuse for interdisciplinary work in Earth and space science disciplines. Networks of diverse people with expertise across Earth, space, and data science disciplines are essential for efficient and ethical exchanges of findable, accessible, interoperable, and reusable (FAIR) research products and practices. Our challenge is then to coordinate the development of standards, curation practices, and tools that enable integrating and reusing multiple data types, software, multi‐scale models, and machine learning approaches across disciplines in a way that is as open and/or FAIR as ethically possible. This is a major endeavor that could greatly increase the pace and potential of interdisciplinary scientific discovery.

58 GEOSCIENCES↗