Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data democratization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

In search of cybernautics

This is a talk about the future of aviation in the information age. Ages come and go. Certainly the atomic age came and went, but the information age looks different. This talk reviews some recent experiments on navigation and control with the Global Positioning System. Vertical position accuracies within 1 foot have been demonstrated in the most recent experiments, and research emphases have shifted to issues of integrity, continuity, and availability. Inertial navigation systems (INS) contribute much to the reliability of GPS-based autoland systems. The GPS data stream can cease, and INS can still complete a precision landing from an altitude of 200 feet. The future of aviation looks like automatic airplanes communicating among each other to schedule ground assets and to avoid collisions and wake hazards. The business of the FAA will be to assure integrity of global navigation systems, to develop and maintain the software rules of the air, and to provide expert pilots to handle emergencies from the ground via radio control. The future of aviation is democratic and lends itself to personal airplanes. Some data analyses reveal that personal airplanes are just as efficient as large turbofan transports and just as fast over distances up to 1,000 miles, thanks to the decelerative influence of the hub and spoke system. Maybe by the year 2020, the airplane will rank with the automobile and computer as an agent of personal freedom.

Crow, Steven

Real-Time Science Decisioning During High Tempo-High Intensity Mission Operations and the Role of Analogs

Introduction: NASA’s VIPER mission presents a unique operational paradigm within the history of robotic spaceflight. The proximity of the Moon to the Earth and the terrain elements (surface characteristics, light/shadow dynamics, communication links) of the lunar South Polar landing site create unprecedented operational conditions between these two planetary bodies. Apollo era lunar science and exploration included humans in situ to operate instruments and assimilate observational inputs in real-time. Previous lunar orbital missions have worked to operational timescales, e.g., decisional timelines and communication exchanges, that were weeks in length. Mars rover missions have worked to operational timescales, e.g., decisional timelines and communication exchanges between Mars and Earth, that were hours, days, and weeks in length. In the case of the VIPER mission, our operational decisioning for rover driving and instrument commanding will be compressed to minute-scale timeframes. These operational conditions directly impact the manner and speed with which the VIPER Science Team (VST) is required to synthesize and analyze data and produce timely science-driven decisions throughout surface mission operations. The VST shall provide mission enhancing scientific input to guide rover traverse planning and drill site confirmation and selection throughout surface operations. Further, the VST input will be of vital importance to the mission’s ability to maximize science return and to meet broader NASA objectives for future lunar in-situ resource utilization (ISRU)and exploration activities. The VST co-located in the Mission Science Center (MSC) will be responsive to the tactical operational cadence of the Mission Operations Center (MOC) and will provide further strategic and Long-Term Planning (LTP) guidance to the mission. The VIPER Science Operations & Integration(SO&I)team has developed an architecture that is focused on the infusion of science-decisioning into the operational framework and execution cadence of VIPER. NASA analog research has played a significant role in the construction of the VIPER science operations systems. As an example, the SO&I team has led analog missions that have focused on bringing together expertise in the sciences (natural, applied and social) and in operations in service of learning how to build and hold together interdisciplinary work environments and what tools are needed to support high tempo, high intensity integrated decisioning. These experiences have provided an essential foundation of knowledge to the VIPER team. Those analogs that specifically influenced the VIPER science operations construct were identified through a process of comparative analysis to prioritize those that offered relevance in whole or in part, and those that did not. The analog research output that provided extensibility to the VIPER science operations architecture included remote teams of humans and robots in cooperation (synchronous and asynchronous) with simulated earthbound systems, engineering and science teams, and the integrated assembly of tools that supported scientific analysis and data synthesis and provided infrastructure for the remote testing framework. Analogs which included real-time data monitoring, synthesis, visualization and access in a democratized and operationalized manner were of particular interest to the development of the VIPER MSC toolset both in terms of the technology and the processes used to develop the supporting infrastructure. We anticipate that each subsequent mission to the lunar south pole, whether with robots or humans, will be able to optimize science and exploration return by evolving strategies to infuse real-time collaborative science-decisioning. Furthermore, these efforts will result in a foundation for science operations development in support of human-robotic exploration of deep space and Mars. NASA analogs can continue to provide the opportunity to prepare, test and iterate on the operational concepts and tools that will support these ever-expanding space exploration efforts. Our presentation will include an overview of the VIPER Science Operations & Integration development process and specifics on what aspects of analog research have had a significant impact on our work systems.

D S S Lim

GeneLab: Omics Database for Spaceflight Experiments

Motivation - To curate and organize expensive spaceflight experiments conducted aboard space stations and maximize the scientific return of investment, while democratizing access to vast amounts of spaceflight related omics data generated from several model organisms. Results - The GeneLab Data System (GLDS) is an open access database containing fully coordinated and curated "omics" (genomics, transcriptomics, proteomics, metabolomics) data, detailed metadata and radiation dosimetry for a variety of model organisms. GLDS is supported by an integrated data system allowing federated search across several public bioinformatics repositories. Archived datasets can be queried using full-text search (e.g., keywords, Boolean and wildcards) and results can be sorted in multifactorial manner using assistive filters. GLDS also provides a collaborative platform built on GenomeSpace for sharing files and analyses with collaborators. It currently houses 172 datasets and supports standard guidelines for submission of datasets, MIAME (for microarray), ENCODE Consortium Guidelines (for RNA-seq) and MIAPE Guidelines (for proteomics).

omics

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]

Ground and Satellite Based Observation Datasets for the Lower Mekong River Basin

In ‘Satellite observations and modeling to understand the Lower Mekong River Basin streamflow variability’ [1] hydrological fluxes, meteorological variables, land cover land use maps, and soil characteristics and parameters data were compiled and processed for the Lower Mekong River Basin. In this work, daily streamflow time series data at nine gauges located at five different countries in the Mekong region (Thailand, Laos People׳s Democratic Republic (PDR), Myanmar, Cambodia, and Viet Nam) is presented. Satellite-based daily precipitation and air temperature (minimum & maximum) data is processed and provided over the entire basin as part of the dataset provided in this work. Moreover, land cover land use raster data that contains 18 classes that cover agriculture, urban, range and forests land cover land use classes for the basin is offered. In addition, a soil data that contains physical and chemical characteristics needed by physically based hydrological models to simulate the cycling of water and air is also provided.

Streamflow

Hacking Kilometer-Scale Models: A Participative Model for Climate Information

In May 2025, nearly 700 participants from all around the world coalesced at 10 regional nodes and a few satellite nodes to take part in a global hackathon of kilometer-scale (horizontal grid spacing < 10 km) regional and global Earth system models. Exciting science is emerging from these efforts, ranging across novel model analysis, new ways of integrating with satellite data, and emulation with machine learning. New technologies were trialed that enable the community to work in new and complementary ways to democratize access to global information at a local scale from a set of the world’s highest-resolution climate models. The hackathon demonstrated how exascale data can be organized to be accessible to anyone. Fundamentally, the community could apply these techniques and technologies to move toward more participative models for coproduction and delivery of diverse sources of climate information for climate scientists and citizens alike.

Climate models

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]

Temporal Explosion Source Processes of Declared Nuclear Tests in the Democratic People’s Republic of Korea

In this work we highlight a preliminary temporal source analysis of the six declared Democratic People's Republic of Korea (DPRK) nuclear tests. We use regional seismic data to estimate relative source time functions (RSTFs) via iterative time-domain deconvolution (Ammon, 2006; Pippin, 2022) of vertical-component ground motions recorded within 2000 km of the source region. Since RSTFs are ideally independent of site and propagation effects, their amplitude spectrum is equivalent to the source spectral ratio, but they also retain phase information. We compare observed RSTFs (in the time and frequency domains) with synthetic RSTFs derived from the Mueller & Murphy (1971) explosion source model. The resolution of these time functions varies, however, we generally obtain high-quality results within the limitations of the recording broadband instrumentation. The results indicate that this method effectively preserves source time-history information that can be used for temporal analysis of remote nuclear explosions. This preliminary analysis is intended to assess the viability of using time-domain deconvolution methods for extracting temporal source information.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Microbial Vessel for Impedance Spectroscopy and Electrochemistry (Mvise): an Extensible, Interoperable Data Acquisition Platform for Liquid Culture Studies in Space Biology Research

The White House Office of Science and Technology Policy (OSTP) has declared 2023 to be the Year of Open Science following an initiative to democratize scientific knowledge. Simultaneously, new sensor technologies have broadened the experimental space available to bioastronautics research. With these open-science goals and technological advances in mind, we have designed and constructed a data acquisition platform for high-precision, real-time monitoring of liquid culture systems. The vessel rig is fitted with six Atlas Scientific probes (micro pH, electrical conductivity, dissolved oxygen, oxidation-reduction potential, liquid temperature, air CO2) and a custom optical density probe similar to the one on BioSentinel’s BioSensor payload. A custom dielectric spectroscopy probe is also planned. The structure of the vessel is resin 3-D printed on a hobbyist-level machine, reducing the production cost and iteration time by over 60% each while increasing extensibility. Data acquisition and storage is controlled with a standalone C state machine-based program running on a Raspberry Pi 3 Model B. When not running headless, an additional program automatically generates and updates plots for live data visualization. Validation of the rig as a data collection system was performed with a yeast liquid culture experiment. While the vessel rig is currently used for standalone experiments, it can also be used as the base perception unit in a self-driving laboratory (SDL). SDLs are high-throughput data collection systems that employ automation and artificial intelligence to conduct and manage routine experiments. Here, we envision an SDL driven by several vessel rigs in which an automated script compares key results, informing the design of future experiments. A vessel rig SDL would streamline many operations, including 1) strain selection for the Lunar Explorer Instrument for space biology Applications (LEIA) investigation and 2) the study of bioregenerative life support systems (BLSS). Ultimately, the datasets that can now be acquired will provide crucial information for accelerating bioastronautics application development in the era of commercial space.

Stephen Lantin

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology

Applying the FAIR Principles to computational workflows

Recent trends within computational and data sciences show an increasing recognition and adoption of computational workflows as tools for productivity and reproducibility that also democratize access to platforms and processing know-how. As digital objects to be shared, discovered, and reused, computational workflows benefit from the FAIR principles, which stand for Findable, Accessible, Interoperable, and Reusable. The Workflows Community Initiative’s FAIR Workflows Working Group (WCI-FW), a global and open community of researchers and developers working with computational workflows across disciplines and domains, has systematically addressed the application of both FAIR data and software principles to computational workflows. We present recommendations with commentary that reflects our discussions and justifies our choices and adaptations. These are offered to workflow users and authors, workflow management system developers, and providers of workflow services as guidelines for adoption and fodder for discussion. The FAIR recommendations for workflows that we propose in this paper will maximize their value as research assets and facilitate their adoption by the wider community.

97 MATHEMATICS AND COMPUTING

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]

Geodetic results from ISAGEX data

Laser and camera data taken during the International Satellite Geodesy Experiment (ISAGEX) were used in dynamical solutions to obtain center-of-mass coordinates for the Astro-Soviet camera sites at Helwan, Egypt, and Oulan Bator, Mongolia, as well as the East European camera sites at Potsdam, German Democratic Republic, and Ondrejov, Czechoslovakia. The results are accurate to about 20m in each coordinate. The orbit of PEOLE (i=15) was also determined from ISAGEX data. Mean Kepler elements suitable for geodynamic investigations are presented.

Marsh, J. G.

CONFLUX: A standardized framework to calculate reactor antineutrino flux

Nuclear fission reactors are abundant sources of antineutrinos for neutrino physics experiments. The flux and spectrum of antineutrinos emitted by a reactor can indicate its activity and composition, suggesting potential applications of neutrino measurements beyond fundamental scientific studies that may be valuable to society. The utility of reactor antineutrinos for applications and fundamental science is dependent on the availability of precise predictions of these emissions. For example, in the last decade, disagreements between reactor antineutrino measurements and models have inspired revision of reactor antineutrino calculations and standard nuclear databases as well as searches for new fundamental particles not predicted by the Standard Model of particle physics. Past predictions and descriptions of the methods used to generate them are documented to varying degrees in the literature, with different modeling teams incorporating a range of methods, input data, and assumptions. The resulting difficulty in accessing or reproducing past models and reconciling results from differing approaches complicates the future study and application of reactor antineutrinos. The CONFLUX (Calculation Of Neutrino FLUX) software framework is a neutrino prediction tool built with the goal of simplifying, standardizing, and democratizing the process of reactor antineutrino flux calculations. CONFLUX includes three primary methods for calculating the antineutrino emissions of nuclear reactors or individual beta decays that incorporate common nuclear data and beta decay theory. The software is prepackaged with the current nuclear databases, including ENDF.B/VIII, JEFF-3.3, and ENSDF, and it includes the capability to predict time-dependent reactor emissions, adjust nuclear database or beta decay inputs/assumptions, and propagate related sources of uncertainty. Here, this paper describes the CONFLUX software structure, details the methods used for flux and spectrum calculations, and provides examples of potential use cases.

Zhang, Xianyi [Lawrence Livermore National Laborat

Critical minerals lists for low-carbon transitions: Reviewing their structure, objectives, and limitations

Critical minerals lists have flourished in the past decade, in particular linked to the importance of critical minerals for low-carbon transitions. We identified 27 critical minerals or materials lists across 15 countries and the European Union (EU). These lists are designed to attract public and private attention and investments to secure both domestic and foreign supplies. This review article fills a gap in the existing literature by analyzing the ways in which these lists are defined and utilized by countries engaged in a mineral rush. We focus our attention on three categories of minerals – battery minerals, platinum-group metals (PGMs), and rare earth elements (REEs) that are particularly important to energy transitions. We situate this research in the broader legal and administrative developments that have driven critical minerals policies in the past decade. We provide an in-depth analysis of the commonalities and variations in the raw materials included in these lists, and identify six core limitations of critical minerals lists: (1) unclear links between criticality assessments and mineral prioritization (2) failure to account for the full mineral value-chain; (3) limited strategic alignment between allied nations; (4) limited flexibility in dynamic environments (5) limited consideration for recycling and by-product sourcing; and (6) reliance on incomplete reserve and resource data.

Battery minerals

Opinion: Aerosol Remote Sensing Over the Next Twenty Years

More than two decades ago, aerosol remote sensing underwent a revolution with the launch of the Terra and Aqua satellites. Advancement continued via additional launches carrying new passive and active sensors. Capable of retrieving parameters characterizing aerosol loading, rudimentary particle properties and in some cases aerosol layer height, the satellite view of Earth’s aerosol system came into focus.The modeling communities have made similar advances. Now the efforts have continued long enough that we can see developing trends in both remote sensing and modeling communities, allowing us to speculate about the future and how the community will approach aerosol remote sensing twenty years from now. We anticipate technology that will replace today’s standard multi-wavelength radiometers with hyperspectral and/or polarimetry all viewing in multiple angles.These will be supported by advanced active sensors with the ability to measure profiles of aerosol extinction in addition to backscatter. The result will be greater insight into aerosol particle properties. Algorithms will move from being primarily physically-based to include an increasing degree of Machine Learning methods, but physically-based techniques will not go extinct. However, the practice of applying algorithms to a single sensor will be in decline. Retrieval algorithms will encompass multiple sensors and all available ground measurements into a unifying framework, and these inverted products will be ingested directly into assimilation systems, becoming “cyborgs”: half observations, half model. In twenty years we will see a true democratization in space with nations large and small, private organizations and commercial entities of all sizes launching space sensors. With this increasing amount of data and aerosol products available, there will be a lot of bad data. User communities will organize to set standards and the large national space agencies will lead the effort to maintain quality by deploying and maintaining validation ground networks and focused field experiments. Through it all, interest will remain high in the global aerosol system and how that system affects climate, clouds, precipitation and dynamics, air quality, the environment and public health, transport of pathogens and fertilization of ecosystems, and how these processes are adapting to a changing climate.

aerosol remote sensing

A Grassroots Network and Community Roadmap for Interconnected Autonomous Science Laboratories for Accelerated Discovery

Scientific discovery is being revolutionized by AI and autonomous systems, yet current autonomous laboratories remain isolated islands unable to collaborate across institutions. We present the Autonomous Interconnected Science Lab Ecosystem (AISLE), a grassroots network transforming fragmented capabilities into a unified system that shorten the path from ideation to innovation to impact and accelerates discovery from decades to months. AISLE addresses five critical dimensions: (1) cross-institutional equipment orchestration, (2) intelligent data management with FAIR compliance, (3) AI-agent driven orchestration grounded in scientific principles, (4) interoperable agent communication interfaces, and (5) AI/ML-integrated scientific education. By connecting autonomous agents across institutional boundaries, autonomous science can unlock research spaces inaccessible to traditional approaches while democratizing cutting-edge technologies. This paradigm shift toward collaborative autonomous science promises breakthroughs in sustainable energy, materials development, and public health.

Ferreira da Silva, Rafael [Oak Ridge National Labo

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection