Engineering PapersSearch

SEARCH · Engineering Papers

Results for “open data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The Pierre Auger Observatory open data

The Pierre Auger Collaboration has embraced the concept of open access to their research data since its foundation, with the aim of giving access to the widest possible community. A gradual process of release began as early as 2007 when 1% of the cosmic-ray data was made public, along with 100% of the space-weather information. In February 2021, a portal was released containing 10% of cosmic-ray data collected by the Pierre Auger Observatory from 2004 to 2018, during the first phase of operation of the Observatory. The Open Data Portal includes detailed documentation about the detection and reconstruction procedures, analysis codes that can be easily used and modified and, additionally, visualization tools. Since then, the Portal has been updated and extended. In 2023, a catalog of the highest-energy cosmic-ray events examined in depth has been included. A specific section dedicated to educational use has been developed with the expectation that these data will be explored by a wide and diverse community, including professional and citizen scientists, and used for educational and outreach initiatives. This paper describes the context, the spirit, and the technical implementation of the release of data by the largest cosmic-ray detector ever built and anticipates its future developments.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

MINERvA s Open Data Product: A First for Neutrino Data Preservation

Access to information on neutrino nucleus interactions is critical to the success of all neutrino oscillation experiments. MINERvA's rich dataset covers a range of energies and nuclei unique amongst experiments, and as such is critical to the community in building the important shared knowledge needed to unravel the mysteries of the neutrino. In particular, its dataset provides the greatest statistical coverage in in the range of neutrino energies pertinent for DUNE until DUNE's near detector begins operation. Historically, such significant datasets in neutrino physics have been preserved primarily through their published results. While meaningful and useful, this limits the ability to explore the data to its fullest extent as new perspectives continue to form. MINERvA has undertaken a major effort to break this trend and preserve its data in a format to be as analyzable as possible from outside the collaboration. This has culminated in the officially-released MINERvA Open Data Product for the community to take advantage of and utilize. Maintaining direct access to the dataset in an analyzable form will allow new insights to continue to be extracted indefinitely. This talk will cover the contents of this product, the information included (and excluded), the tools provided to utilize the product effectively, the support MINERvA intends to provide in its use, and some lessons learned through the process.

Last, David [Rochester U.] (ORCID:0000000245147183

The UARS and open data system concept and analysis study. Executive summary

Alternative concepts for a common design for the UARS and OPEN Central Data Handling Facility (CDHF) are offered. The designs are consistent with requirements shared by UARS and OPEN and the data storage and data processing demands of these missions. Because more detailed information is available for UARS, the design approach was to size the system and to select components for a UARS CDHF, but in a manner that does not optimize the CDHF at the expense of OPEN. Costs for alternative implementations of the UARS designs are presented showing that the system design does not restrict the implementation to a single manufacturer. Processing demands on the alternative UARS CDHF implementations are discussed. With this information at hand together with estimates for OPEN processing demands, it is shown that any shortfall in system capability for OPEN support can be remedied by either component upgrades or array processing attachments rather than a system redesign.

Mittal, M.

The UARS and open data concept and analysis study

Alternative concepts for a common design for the UARS and OPEN Central Data Handling Facility (CDHF) are offered. Costs for alternative implementations of the UARS designs are presented, showing that the system design does not restrict the implementation to a single manufacturer. Processing demands on the alternative UARS CDHF implementations are then discussed. With this information at hand together with estimates for OPEN processing demands, it is shown that any shortfall in system capability for OPEN support can be remedied by either component upgrades or array processing attachments rather than a system redesign. In addition to a common system design, it is shown that there is significant potential for common software design, especially in the areas of data management software and non-user-unique production software. Archiving the CDHF data are discussed. Following that, cost examples for several modes of communications between the CDHF and Remote User Facilities are presented. Technology application is discussed.

Mittal, M.

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY

Open Data Integration (ODIN): A Concurrent, Distributed Message-Based Architecture and Framework for Disaster Response

The Runtime for Airspace Concept Evaluation (RACE) is an open-source software architecture and framework to build configurable, highly concurrent and distributed message-based systems that offer scalable, low-latency performance on commodity hardware. RACE was used in commercial aviation applications to rapidly build systems that span several machines (including synchronized displays), interface existing hardware simulators and other live data feeds, and incorporate sophisticated visualization components such as NASA WorldWind. These RACE applications validated elements of the FAA’s System Wide Information Management (SWIM) Program, handling up to 1000 messages/sec from diverse sources (SFDPS, TFM-DATA, TAIS, ASDE-X, ITWS and local ADS) for 4,500 simultaneous flights tracked in the next-generation air transportation system’s digital backbone. We have since generalized RACE to support Open Data Integration (ODIN) applications outside aviation. Systems built with RACE/ODIN can be deployed in the field, on commodity hardware, and operate with limited or intermittent connectivity to the outside world. Our primary use case is a web-server with local/persistent data storage that runs within and only serves the stakeholder network (e.g. an incident command post). We are tailoring the RACE/ODIN system to support wildland fire management for the upcoming NASA Wildland Fire Safety Demonstration Series. RACE-ODIN is under consideration for application in the Scalable Traffic Management for Emergency Response Operations project, or STEReO, which aims to create a system that can be deployed during emergencies, to coordinate multiple elements of disaster response. Such data sources predominantly come from existing services on the internet (e.g. weather and satellite data, imported from so called "edge servers") but can also include dynamic (real-time) data from computer simulations and within the stakeholder network (such as aircraft and personnel tracking information). We will present the architecture and ODIN system demonstration incorporating local data from instrumented power-line towers, interpolated weather data and geospatial data from space-based platforms.

Joseph C Coughlan

Open Data Integration (ODIN): A Concurrent, Distributed Message-Based Architecture and Framework for Disaster Response

The Runtime for Airspace Concept Evaluation (RACE) is an open-source software architecture and framework to build configurable, highly concurrent and distributed message-based systems that offer scalable, low-latency performance on commodity hardware. RACE was used in commercial aviation applications to rapidly build systems that span several machines (including synchronized displays), interface existing hardware simulators and other live data feeds, and incorporate sophisticated visualization components such as NASA WorldWind. These RACE applications validated elements of the FAA’s System Wide Information Management (SWIM) Program, handling up to 1000 messages/sec from diverse sources (SFDPS, TFM-DATA, TAIS, ASDE-X, ITWS and local ADS) for 4,500 simultaneous flights tracked in the next-generation air transportation system’s digital backbone. We have since generalized RACE to support Open Data Integration (ODIN) applications outside aviation. Systems built with RACE/ODIN can be deployed in the field, on commodity hardware, and operate with limited or intermittent connectivity to the outside world. Our primary use case is a web-server with local/persistent data storage that runs within and only serves the stakeholder network (e.g. an incident command post). We are tailoring the RACE/ODIN system to support wildland fire management for the upcoming NASA Wildland Fire Safety Demonstration Series. RACE-ODIN is under consideration for application in the Scalable Traffic Management for Emergency Response Operations project, or STEReO, which aims to create a system that can be deployed during emergencies, to coordinate multiple elements of disaster response. Such data sources predominantly come from existing services on the internet (e.g. weather and satellite data, imported from so called "edge servers") but can also include dynamic (real-time) data from computer simulations and within the stakeholder network (such as aircraft and personnel tracking information). We will present the architecture and ODIN system demonstration incorporating local data from instrumented power-line towers, interpolated weather data and geospatial data from space-based platforms.

Guillaume P Brat

NASA Open Science Data Repository: Open Science for Life in Space

Space biology and health data are critical for the success of deep space missions and sustainable human presence off-world. At the core of effectively managing biomedical risks is the commitment to open science principles, which ensure that data are findable, accessible, interoperable, reusable, reproducible and maximally open. The 2021 integration of the Ames Life Sciences Data Archive with GeneLab to establish the NASA Open Science Data Repository significantly enhanced access to a wide range of life sciences, biomedical-clinical, and mission telemetry data alongside existing ‘omics data from GeneLab. This paper describes the new database, its architecture, and new data streams supporting diverse data types and enhancing data submission, retrieval, and analysis. Features include the Biological Data Management Environment for improved data submission, a new user interface, controlled data access, an enhanced API, and comprehensive public visualization tools for environmental telemetry, radiation dosimetry data, and ‘omics analyses. By fostering global collaboration through its Analysis Working Groups and training programs, the Open Science Data Repository promotes widespread engagement in space biology, ensuring transparency and inclusivity in research. It supports the global scientific community in advancing our understanding of spaceflight's impact on biological systems, ensuring humans will thrive in future deep space missions.

OSDR

How the Open Data Policy of the Landsat Program Has Advanced Our Understanding of Environmental Change

A time series is a sequence of observations of a phenomenon taken sequentially in time. A crucial characteristic of a time series is the dependence among adjacent observations – techniques for analyzing this dependence are referred to as time series analysis. This analytical approach enables us to predict or forecast future values of a time series, study the impact of various inputs on the observed phenomenon, and examine interrelationships among related time series variables. Within the geographical sciences, time series analysis has historically been limited to coarse-resolution satellite data, as constructing time series of data suitable for studying land cover and land use dynamics, such as Landsat data, were prohibitively costly. A transformative shift occurred in 2008 when the U.S. Government decided to make free and open all past and future data collected by the Landsat satellite program. The decision brought about a paradigm shift away from analyzing individual images or observations to continuous monitoring in time. Of particular relevance to environmental remote sensing is the ability to forecast observations – if we can predict how future observations should behave, we can infer information about how the land surface is changing. In this presentation, we will examine literature examples that showcase scientific gains enabled by time series analysis of satellite data. We will delve into how the analysis of dense time series of satellite data revealed that overall rate of forest disturbance in the Amazon has increased despite a reduction in deforestation; how different types of forest degradation, previously unquantified, are now being accurately assessed in the Caucasus region; and how we now can study the highly dynamic and intricate patterns of shifting cultivation in Southeast Asia.

Pontus Olofsson

High-Resolution Satellite Data Open for Government Research

U.S. satellite commercial imagery (CI) with resolution less than 1 meter is a common geospatial reference used by the public through Web applications, mobile devices, and the news media. However, CI use in the scientific community has not kept pace, even though those who are performing U.S. government research have access to these data at no cost.Previously, studies using multiple CI acquisitions from IKONOS-2, Quickbird-2, GeoEye-1, WorldView-1, and WorldView-2 would have been cost prohibitive. Now, with near-global submeter coverage and online distribution, opportunities abound for future scientific studies. This archive is already quite extensive (examples are shown in Figure 1) and is being used in many novel applications.

Data

An open retail boundary dataset for South Korea using open data and computer vision technique

Although delineating retail boundaries is important to explore and comprehend the dynamics of the retail sector, it is hard to find studies specifically addressing it in the South Korean context. This study fills this gap by proposing new retail boundaries across South Korea. To achieve this goal, we employed a variety of retailers and building datasets and proposed a unique computer vision-based framework with a deep ensemble voting technique. As a result, we delineated 6,636 distinct retail boundaries that were validated against existing reference retail boundaries. These newly delineated retail boundaries provide valuable insights for researchers, governments, and other relevant stakeholders by enhancing their understanding of retail geography. This dataset can be used as a foundational resource for analyses on topics such as pandemic recovery, retail gentrification, and the resilience of retail spaces in response to e-commerce growth, ultimately contributing to more robust retail sector research in South Korea.

97 MATHEMATICS AND COMPUTING

Filling in Subsurface Storage Open Data Gaps - Updates to CCS Data Availability on EDX and EDX Spatial (FWP-1022465)

There is a need to preserve and efficiently access data resources to drive the next generation of research and development while ensuring compliance with DOE regulations. Over the last 10+ years, there has been ongoing efforts by the DOE Carbon Storage Program to ensure that there is effective data curation and preservation of DOE funded research leveraging the NETL-FECM data repository, the Energy Data eXchange (EDX). This talk presents updates about ongoing efforts to continue to support the mission of ensuring that carbon storage data is findable, accessible, interoperable, and reusable to the carbon storage stakeholder community through EDX and EDX Spatial. Presented at the NETL Carbon Management Review Meeting, Pittsburgh, 2024.

Morkner, Paige

MINERvA open-data product

MINERvA is THE neutrino cross section experiment Scintillator tracker/calorimeter ran in the NuMI beam at Fermilab same beam as the MINOS and NOvA oscillation experiments With our data, we are solving systematic shortcomings in neutrino interaction rate/spectra that are the largest part of the systematic uncertainty in today s (and tomorrow s) measurements. Some aspects have NO equivalent in the neutrino program future. Scientific scope includes both particle and nuclear physics GeV scale cross sections, A dependence, MeV scale effects most published measurements for a neutrino experiments ever (well, tied with T2K) with 25% more papers in the pipeline.

Gran, Rik [Minnesota U., Duluth] (ORCID:0000000216

Sharing is Caring: A Practical Guide to FAIR(ER) Open Data Release

This is a two hour version of the FAIR(ER) tutorial we released at Barcelona 9/24. SAND2024-12152C. The only modifications were largely deletions, which don't require additional review. The one key difference that actually has changed material is in the Language section for Equitable Accessibility, which is almost word for word the same as previously approved SAND2025-04087W which is the website version of the presentation.

Henriksen, Amelia [Sandia National Laboratories (S

Linked Open Data in the Global Change Information System (GCIS)

The U.S. Global Change Research Program (http://globalchange.gov) coordinates and integrates federal research on changes in the global environment and their implications for society. The USGCRP is developing a Global Change Information System (GCIS) that will centralize access to data and information related to global change across the U.S. federal government. The first implementation will focus on the 2013 National Climate Assessment (NCA) . (http://assessment.globalchange.gov) The NCA integrates, evaluates, and interprets the findings of the USGCRP; analyzes the effects of global change on the natural environment, agriculture, energy production and use, land and water resources, transportation, human health and welfare, human social systems, and biological diversity; and analyzes current trends in global change, both human-induced and natural, and projects major trends for the subsequent 25 to 100 years. The NCA has received over 500 distinct technical inputs to the process, many of which are reports distilling and synthesizing even more information, coming from thousands of individuals around the federal, state and local governments, academic institutions and non-governmental organizations. The GCIS will present a web-based version of the NCA including annotations linking the findings and content of the NCA with the scientific research, datasets, models, observations, etc. that led to its conclusions. It will use semantic tagging and a linked data approach, assigning globally unique, persistent, resolvable identifiers to all of the related entities and capturing and presenting the relationships between them, both internally and referencing out to other linked data sources and back to agency data centers. The developing W3C PROV Data Model and ontology will be used to capture the provenance trail and present it in both human readable web pages and machine readable formats such as RDF and SPARQL. This will improve visibility into the assessment process, increase understanding and reproducibility, and ultimately increase credibility and trust of the resulting report. Building on the foundation of the NCA, longer term plans for the GCIS include extending these capabilities throughout the U.S. Global Change Research Program, centralizing access to global change data and information across the thirteen agencies that comprise the program.

Tilmes, Curt A.

The Open Data Repository's Data Publisher

We have quickly gathered a diverse set of databases that have significant activity. Researchers are using them on a daily basis from collection through all phases of research. Focus on: (1) Linked data and semantic web integration are fundamental to our long term plans. (2) Citation system is critical before our first full release. Snapshot-based citation generation for data sets or objects. As we expand our feature set, expect the system will: (1) Provide a very high level of provenance for any data housed in the software. (2) Make data management beneficial to the researcher throughout the research process. (3) Create living archives that allow citable snapshots of data and summaries of the changes since the snapshot was created.

Stone, N.