Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDR's success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

access↗

Utilizing AI and Spatial Data to Identify & Rapidly Disseminate Energy Infrastructure Insights

GeoGov Summit Final Presentation entitled "Utilizing AI and Spatial Data to Identify & Rapidly Disseminate Energy Infrastructure Insights". Maintaining the integrity of energy infrastructure plays a critical role in ensuring energy security. Robust foundational AI models using data from federal, state, industry, and other sources can help address integrity risk management & mitigation issues as well as evaluate extended use strategies. Trusted foundational models can help with industry adoption and accelerate innovation by enhancing integrity predictions, reduce costs, and informing infrastructure build-out. Coordination, collaboration & data sharing to develop robust models to aid in: Optimizing operations; Minimizing costs; Ensuring energy security.

Advanced Infrastructure Integrity Model (AIIM)↗

Modeling of Reactor Design and Optimization for Scale-Up of the Catalyxx Process for Ethanol Conversion to Higher Alcohol Biofuels

This report summarizes the results of a collaborative efforts between Oak Ridge National Laboratory (ORNL) and Catalyxx Inc. to investigate scale-up of Catalyxx’s Ethanol upgrading to higher alcohols process. The study was funded by the U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) under CRADA (Cooperative Research and Development Agreement) No: NFE-20-08396. The project is part of the Direct Funding Opportunity (DFO) for Computational Science to Enable Bioenergy program which utilized computational toolsets developed by the Consortium for Computational Physics and Chemistry, a multi-laboratory consortium in BETO. The report here summarizes a packed-bed reactor modeling effort spanning the range from lab to industrial scale (from 4 gram to 5-ton catalyst beds), and examining reactor design, process optimization strategies, and suggested design and operating conditions for Catalyxx’s ethanol upgrading plants. The results in this report have been shared in monthly steering meetings and presentations are available in the shared data house owned by Catalyxx. The modeling effort helped to define optimum operation conditions for maximum alcohol selectivity and yield: i.e., temperature control scenarios ranging from adiabatic to isothermal, feed rate, pressure, and inlet H 2 /Ethanol ratio. The modeling results were verified at lab-(4 gram) and pre-pilot (4 kg) scales and has been used to evaluate a 5-ton packed-bed reactor and identify operating conditions to maximize the butanol yield. Special focus was given to understanding mass-transfer effects in the pre-pilot and pilot-scale reactors, over the domain of flow rate, pressure, feed composition, pellet size, shape, porosity, bed voidage, and reactor dimensions (i.e., length/diameter). Modeling was also used to evaluate innovative reactor design concepts such as water removal to improve alcohol selectivity and yield, and a reactor with an additional side inlet to facilitate quenching. These concepts were thoroughly explored, and potential benefits were disclosed. The results in this report are summarized and described qualitatively to protect the IP rights of Catalyxx. The details have been shared with the Catalyxx team in the regular steering meetings. At the end of the project, Catalyxx Inc. announced a successful demonstration of pilot scale operation in Seville, Spain.

02 PETROLEUM↗

Exponential Backoff and Its Security Implications for Safety-Critical OT Protocols over TCP/IP Networks

The convergence of Operational Technology (OT) and Information Technology (IT) networks has become increasingly prevalent with the growth of Industrial Internet of Things (IIoT) applications. This shift, while enabling enhanced automation, remote monitoring, and data sharing, also introduces new challenges related to communication latency and cybersecurity. Oftentimes, legacy OT protocols were adapted to the TCP/IP stack without an extensive review of the ramifications to their robustness, performance, or safety objectives. To further accommodate the IT/OT convergence, protocol gateways were introduced to facilitate the migration from serial protocols to TCP/IP protocol stacks within modern IT/OT infrastructure. However, they often introduce additional vulnerabilities by exposing traditionally isolated protocols to external threats. This study investigates the security and reliability implications of migrating serial protocols to TCP/IP stacks and the impact of protocol gateways, utilizing two widely used OT protocols: Modbus TCP and DNP3. Our protocol analysis finds a significant safety-critical vulnerability resulting from this migration, and our subsequent tests clearly demonstrate its presence and impact. A multi-tiered testbed, consisting of both physical and emulated components, is used to evaluate protocol performance and the effects of device-specific implementation flaws. Through this analysis of specifications and behaviors during communication interruptions, we identify critical differences in fault handling and the impact on time-sensitive data delivery. The findings highlight how reliance on lower-level IT protocols can undermine OT system resilience, and they inform the development of mitigation strategies to enhance the robustness of industrial communication networks.

DNP3↗

Specifying Calibration of Environmental Sensors

The emergence of the Internet of Things is resulting in an increased ability of devices and systems to share data and is generating increasing interest in integrating sensors into a variety of devices deployed in the built environment. The value of such data is a function of how the data can be used. Data-producing devices and systems that enable valuable use-cases in turn can be seen as more valuable. Lighting systems are particularly interesting platforms for integrated sensors. Both indoor and outdoor lighting devices are becoming more connected, and their location is often ideally suited for hosting environmental sensors that can characterize the properties of indoor or outdoor spaces in ways that support a wide variety of use cases, from improving air quality to supporting fault diagnostics and prediction. The value of environmental-sensor-driven use-cases and the lighting systems that house them is dependent to some degree on sensor accuracy. Environmental sensors utilize a wide variety of sensing techniques or technologies and have varying accuracy. More-accurate, laboratory-grade products or reference standards are often used to characterize, refine, calibrate, adjust, and monitor devices that are deployed, or are intended to be deployed, in physical spaces of interest. Sensors or reference standards need to be calibrated periodically to ensure that their use yields accurate measurements. Calibration needs, however, vary in sophistication, based on user and use-case requirements. This paper provides guidance for evaluating the performance of environmental sensors so as to ensure that they meet user or use-case needs. It describes best practices that have been developed for a) calibrating sensors to ensure some known level of accuracy, and b) determining whether calibration-laboratory accreditation meets user or use-case needs. Excerpts from laboratory scopes of accreditation are shared to reveal the diversity of terminology and format among them. In an effort to aid those who currently have sensors calibrated or who have new or changing needs for sensor calibration, rationale is provided for why a specification might be used to request calibration services that meet specific needs. Commercially available calibration-service providers that are accredited for environmental-sensor calibration are compared and contrasted, and a specification template that might be used for requesting this calibration is presented. The specification template should be tailored to meet each user’s needs. To illustrate, an example set of environmental-sensor test conditions (reflecting the planned usage of the device to be calibrated) is used to develop a customized calibration specification, and commercially available service providers are assessed in terms of their qualifications for calibration to that particular implementation of the specification template.

42 ENGINEERING↗

Hawaii Play Fairway Analysis: Noble Gas Raw Data for Hawaii, Maui, Oahu, Kauai, and Lanai islands

Noble gas raw data for the Hawaiian islands of Big Island, Maui, Oahu, Lanai, and Kauai. Based on results from prior phases of the Hawaii Play Fairway Analysis, this project targeted 66 wells on the islands of Hawaii, Maui, Lanai, Oahu, and Kauai for sampling of dissolved noble gases, trace metals, common ions, and the stable isotopes 2H and 18O. Ultimately, 23 of the 66 well targets were sampled. Noble gas data from this study is supplemented with data shared by the United States Geologic Survey for the summit of Kilauea, and by the geothermal energy company Ormat Technologies Inc. for their geothermal power plant Puna Geothermal Venture on the Lower East Rift of Kilauea, and for their exploration of Kona and Hualalai on Hawaii, as well as the Southwest Rift of Haleakala on Maui. The noble gas helium is used as an indicator of geothermal heat when excess 3He and/or 4He is present when compared to the atmospheric ratio of those isotopes (R/Ra). R/Ra is minimally affected by dilution and transport, allowing even those wells not perfectly situated over a geothermal system to indicate a geothermal anomaly. R/Ra anomalies are present on every island in this study. There is a strong correlation between R/Ra anomalies and proximity to rift zones and calderas. The Hawaii Play Fairway project was funded by the U.S. Department of Energy Geothermal Technologies Office (award DE-EE0006729). For more information, see Colin Ferguson's Master of Science thesis "Exploration for Blind Geothermal Resources in the State of Hawaii Utilizing Dissolved Noble Gasses in Well Waters."

15 GEOTHERMAL ENERGY↗

Position-Enhanced Gradient Attack (PEGA) on Medical Language Models

Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.

Xu, Nuo [University of Minnesota]↗

SAM-I-Am: Semantic boosting for zero-shot atomic-scale electron micrograph segmentation

Image segmentation is a critical enabler for tasks ranging from medical diagnostics to autonomous driving. However, the correct segmentation semantics — where are boundaries located? what segments are logically similar? — change depending on the domain, such that state-of-the-art foundation models can generate meaningless and incorrect results. Moreover, in certain domains, fine-tuning and retraining techniques are infeasible: obtaining labels is costly and time-consuming; domain images (micrographs) can be exponentially diverse; and data sharing (for third-party retraining) is restricted. To enable rapid adaptation of the best segmentation technology, we propose the concept of semantic boosting: given a zero-shot foundation model, guide its segmentation and adjust results to match domain expectations. Here, we apply semantic boosting to the Segment Anything Model (SAM) to obtain microstructure segmentation for transmission electron microscopy. Our booster, SAM-I-Am, serves as a post-processing engine that extracts geometric and textural features of various intermediate masks to perform mask removal and mask merging operations. We demonstrate a zero-shot performance increase of (absolute) +21.35%, +12.6%, +5.27% in mean IoU, and a -9.91%, -18.42%, -4.06% drop in mean false positive masks across images of three difficulty classes over vanilla SAM (ViT-L).

36 MATERIALS SCIENCE↗

A Phage Foundry Framework to Systematically Develop Viral Countermeasures to Combat Antibiotic-Resistant Bacterial Pathogens

At its current rate, the rise of antimicrobial-resistant (AMR) infections is predicted to paralyze our industries and healthcare facilities while becoming the leading global cause of loss of human life. With limited new antibiotics on the horizon, we need to invest in alternative solutions. Bacteriophages (phages)–viruses targeting bacteria–offer a powerful alternative approach to tackle bacterial infections. Despite recent advances in using phages to treat recalcitrant AMR infections, the field lacks systematic development of phage therapies scalable to different applications. We propose a Phage Foundry framework to establish metrics for phage characterization and to fill the knowledge and technological gaps in phage therapeutics. Coordinated investment in AMR surveillance, sampling, characterization, and data sharing procedures will enable rational exploitation of phages for treatments. A fully realized Phage Foundry will enhance the sharing of knowledge, technology, and viral reagents in an equitable manner and will accelerate the biobased economy.

59 BASIC BIOLOGICAL SCIENCES↗

The Data Synergy Effects of Time-Series Deep Learning Models in Hydrology

When fitting statistical models to variables in geoscientific disciplines such as hydrology, it is a customary practice to stratify a large domain into multiple regions (or regimes) and study each region separately. Traditional wisdom suggests that models built for each region separately will have higher performance because of homogeneity within each region. However, each stratified model has access to fewer and less diverse data points. Here, through two hydrologic examples (soil moisture and streamflow), we show that conventional wisdom may no longer hold in the era of big data and deep learning (DL). We systematically examined an effect we call data synergy, where the results of the DL models improved when data were pooled together from characteristically different regions. The performance of the DL models benefited from modest diversity in the training data compared to a homogeneous training set, even with similar data quantity. Moreover, allowing heterogeneous training data makes eligible much larger training datasets, which is an inherent advantage of DL. A large, diverse data set is advantageous in terms of representing extreme events and future scenarios, which has strong implications for climate change impact assessment. The results here suggest the research community should place greater emphasis on data sharing.

54 ENVIRONMENTAL SCIENCES↗

A Dive into Underwater Solar Cells

Our oceans are vast, mostly unexplored and difficult to monitor. Large-scale implementation of a fully autonomous 'Internet of Underwater Things' would transform how we collect and share data from this domain; however, deployment is prohibited by the lack of persistent power sources. In principle, underwater solar-energy generation can complement the use of batteries and provide a solution, although dedicated research is needed since traditional silicon solar cells do not perform well underwater due to water's strong absorption of near-infrared light. In this Perspective we present examples of solar-powered underwater applications and discuss which types of solar-harvesting materials could be appropriate, including GaInP variants, CdTe, organic semiconductors, and perovskite semiconductors. We also discuss challenges that need to be addressed, such as the development of effective antifouling coatings and new certification standards given that underwater conditions are starkly different from those in terrestrial environments.

antifouling coatings↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

White paper on light sterile neutrino searches and related phenomenology

This white paper provides a comprehensive review of our present understanding of experimental neutrino anomalies that remain unresolved, charting the progress achieved over the last decade at the experimental and phenomenological level, and sets the stage for future programmatic prospects in addressing those anomalies. It is purposed to serve as a guiding and motivational "encyclopedic" reference, with emphasis on needs and options for future exploration that may lead to the ultimate resolution of the anomalies. We see the main experimental, analysis, and theory-driven thrusts that will be essential to achieving this goal being: 1) Cover all anomaly sectors -- given the unresolved nature of all four canonical anomalies, it is imperative to support all pillars of a diverse experimental portfolio, source, reactor, decay-at-rest, decay-in-flight, and other methods/sources, to provide complementary probes of and increased precision for new physics explanations; 2) Pursue diverse signatures -- it is imperative that experiments make design and analysis choices that maximize sensitivity to as broad an array of these potential new physics signatures as possible; 3) Deepen theoretical engagement -- priority in the theory community should be placed on development of standard and beyond standard models relevant to all four short-baseline anomalies and the development of tools for efficient tests of these models with existing and future experimental datasets; 4) Openly share data -- Fluid communication between the experimental and theory communities will be required, which implies that both experimental data releases and theoretical calculations should be publicly available; and 5) Apply robust analysis techniques -- Appropriate statistical treatment is crucial to assess the compatibility of data sets within the context of any given model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A use-case-driven approach for demonstrating the added value of digitalisation in wind energy

Digitalisation is one of the key drivers for reducing the costs and risks of wind energy. When considering whether to embark on a digitalisation initiative, two key questions arise. The first is what business or operational opportunities might feasibly be addressed and the second is which of the many potential aspects of digitalisation are relevant to those opportunities. In this work, we show how these questions can be answered with a use-case-driven approach, based around a survey aiming to collect and collate the main "pain points" (or everyday challenges) of people in the wind energy sector. Although the relatively low number of participants of the survey (46) means that the results should only be used indicatively, it is still possible to make some general recommendations for priorities for digitalisation efforts in the wind energy sector. Firstly, digitalisation efforts should focus both on supporting people carrying out cross-lifecycle tasks, in particular sharing data, managing data, undertaking general data analyses and accessing data. Tools to do this should deal with varying data formats and naming conventions, make metadata more accessible, define data and metadata standards, make more data publicly available and improve the quality of data. Secondly, efforts should also focus on supporting people in the wind farm operational phase, in particular with failure detection, fault diagnosis, failure rate modelling and predictive maintenance. Solutions to do this should focus on accessible and validated tools for fault detection, cloud or other data pipeline solutions for SCADA data and tools for exhaustive data documentation. Finally, digitalisation efforts should focus on better communicating and helping people become aware of existing solutions and tools, as well as on helping people to exert a stronger influence on possible solutions.

17 WIND ENERGY↗

BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data

Abstract BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models and computational chemistry methods development. This update reports significant growth and enhancements since our last review in 2016. Of note, the database now contains 2.9 million binding measurements spanning 1.3 million compounds and thousands of protein targets. This growth is largely attributable to our unique focus on curating data from US patents, which has yielded a substantial influx of novel binding data. Recent improvements include a remake of the website following responsive web design principles, enhanced search and filtering capabilities, new data download options and webservices and establishment of a long-term data archive replicated across dispersed sites. We also discuss BindingDB’s positioning relative to related resources, its open data sharing policies, insights gleaned from the dataset and plans for future growth and development.

Liu, Tiqing↗

Frontiers and opportunities in bioenergy crop microbiome research networks

Researchers from across the four U.S. Department of Energy Bioenergy Research Centers engaged in a microbiome workshop that focused on identifying challenges and collaboration opportunities to better understand bioenergy-relevant plant–microbe interactions. The virtual workshop included hands-on educational sessions and a keynote address on current best practices in microbiome science and community microbiome standards, as well as breakout sessions aimed at identifying microbiome-related data and measurements that should be prioritized, opportunities for and barriers to integrating plant metabolites to microbiome research, and strategies for more effectively integrating microbiome data and processes into existing models. Based on participant discussion, key findings of the workshop were the need to prioritize scaling data sharing across BRCs and the broader research community and securing collaborative infrastructure in the areas of microbiome-ecosystem modeling and molecular plant-microbe interactions. This workshop review highlights additional main findings from this event, to encourage cross-site and more holistic meta-analyses while promoting wide scientific community engagement across plant microbiome sciences.

09 BIOMASS FUELS↗

Evaluating Recursive Blind Forecast Against API and Baseline: A Puerto Rican Case Study on Solar Irradiance for Normal and Extreme Weather

This paper leverages ongoing work in a community microgrid in Adjuntas, Puerto Rico to forecast global horizontal irradiance (GHI) and compare performance in normal and extreme weather. Given a positive correlation of 0.98 between GHI and PV power, forecasting GHI can be an effective, indirect forecast of photovoltaic (PV) power, especially in microgrids where the end-users, owners, operators, or other stakeholders are reluctant to share data for training or validation due to privacy and security concerns. A recursive one-shot (termed as "blind") forecast is, hence, formulated, wherein a gradient-boosted regression tree (GBR) is built to forecast GHI for a 7-day horizon in normal weather, and a 2-day horizon in extreme weather. To demonstrate its resilience, the architecture is trained on normal and hurricane weather GHI from 2002-2022. It is generalized on February 9-16, 2023, and on the landfall of Hurricane Nicole (Nov 4-5, 2022), respectively. Forecasts from GBR are compared against that from a satellite-based API resource and three baselines: persistence, averaging, and exponential smoothing. Results show GBR and persistence outperform sophisticated API in both types of weather for this case study.

Sundararajan, Aditya↗

Effectiveness and predictability of in-network storage cache for Scientific Workflows

Large scientific collaborations often have multiple scientists accessing the same set of files while doing different analyses, which create repeated accesses to the large amounts of shared data located far away. These data accesses have long latency due to distance and occupy the limited bandwidth available over the wide-area network. To reduce the wide-area network traffic and the data access latency, regional data storage caches have been installed as a new networking service. To study the effectiveness of such a cache system in scientific applications, we examine the Southern California Petabyte Scale Cache for a high-energy physics experiment. By examining about 3TB of operational logs, we show that this cache removed 67.6% of file requests from the wide-area network and reduced the traffic volume on wide-area network by 12. 3TB (or 35.4%) an average day. The reduction in the traffic volume (35.4%) is less than the reduction in file counts (67.6%) because the larger files are less likely to be reused. Due to this difference in data access patterns, the cache system has implemented a policy to avoid evicting smaller files when processing larger files. We also build a machine learning model to study the predictability of the cache behavior. Tests show that this model is able to accurately predict the cache accesses, cache misses, and network throughput, making the model useful for future studies on resource provisioning and planning.

Sim, Caitlin↗