Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Adapting Secure MultiParty Computation to Support Machine Learning in Radio Frequency Sensor Networks

In this project we developed and validated algorithms for privacy-preserving linear regression using a new variant of Secure Multiparty Computation (MPC) we call "Hybrid MPC" (hMPC). Our variant is intended to support low-power, unreliable networks of sensors with low-communication, fault-tolerant algorithms. In hMPC we do not share training data, even via secret sharing. Thus, agents are responsible for protecting their own local data. Only the machine learning (ML) model is protected with information-theoretic security guarantees against honest-but-curious agents. There are three primary advantages to this approach: (1) after setup, hMPC supports a communication-efficient matrix multiplication primitive, (2) organizations prevented by policy or technology from sharing any of their data can participate as agents in hMPC, and (3) large numbers of low-power agents can participate in hMPC. We have also created an open-source software library named "Cicada" to support hMPC applications with fault-tolerance. The fault-tolerance is important in our applications because the agents are vulnerable to failure or capture. We have demonstrated this capability at Sandia's Autonomy New Mexico laboratory through a simple machine-learning exercise with Raspberry Pi devices capturing and classifying images while flying on four drones.

42 ENGINEERING↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Dynamic Boundary Microgrids Under Privatization Considerations

Microgrids have physical, electrical, and logical (data, network, and ownership) boundaries. To power unserved customer loads during an outage, microgrids can extend the traditional operational boundaries. This can become complex when considering microgrid-to-microgrid (M2M) interactions where sensitive information such as competitive microgrid operational data is not shared. This work proposes an optimization method coordinated between microgrid controllers and distribution management systems that limits data sharing. The method involves a competitive bidding strategy that maximizes unserved load coverage while minimizing resource utilization and sensitive operational data sharing among entities. The work is validated on a two-microgrid system with photovoltaic and energy storage systems and curves of load derived from real world residential buildings datasets. Results show that the proposed method, when applied for three distinct use cases of energy storage sufficiency to cover the predefined boundary and/or the expanded boundary, can successfully select and bid the available load coverage.

Starke, Michael [ORNL] (ORCID:0000000221211195)↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

Cosmological constraints from the Planck cluster catalogue with DES shear profiles and Chandra observations

We present cosmological constraints from the Planck PSZ2 cosmological cluster sample, using weak-lensing shear profiles from Dark Energy Survey (DES) data and X-ray observations from the Chandra telescope for the mass calibration. We compute hydrostatic mass estimates for all clusters in the PSZ2 sample with a scaling relation between their Sunyaev-Zeldovich signal and X-ray derived hydrostatic mass, calibrated with the Chandra data. We introduce a method to correct these masses with a hydrostatic mass bias using shear profiles from wide-field galaxy surveys. We simultaneously fit the number counts of the PSZ2 sample and the mass calibration with the DES data, finding $Ω_\text{m}=0.312^{+0.018}_{-0.024}$, $σ_8=0.777\pm 0.024$, $S_8\equiv σ_8 \sqrt{Ω_\text{m} / 0.3}=0.791^{+0.023}_{-0.021}$, and $(1-b)=0.844^{+0.055}_{-0.062}$ for our baseline analysis when combined with BAO data. When considering a hydrostatic mass bias evolving with mass, we find $Ω_\text{m}=0.353^{+0.025}_{-0.031}$, $σ_8=0.751\pm 0.023$, and $S_8=0.814^{+0.019}_{-0.020}$. We verify the robustness of our results by exploring a variety of analysis settings, with a particular focus on the definition of the halo centre used for the extraction of shear profiles. We compare our results with a number of other analyses, in particular two recent analyses of cluster samples obtained from SPT and eROSITA data that share the same mass calibration data set. We find that our results are in overall agreement with most late-time probes, in very mild tension with CMB results (1.6$σ$), and in significant tension with results from eROSITA clusters (2.9$σ$). We confirm that our mass calibration is consistent with the eROSITA analysis by comparing masses for clusters present in both Planck and eROSITA samples, eliminating it as a potential cause of tension.

Aymerich, G. [Orsay, IAS; AIM, Saclay] (ORCID:0009↗

Livewire Data Platform: A Catalog of Transportation and Mobility Data

A two-page fact sheet on the capabilities and core services of the Livewire Data Platform, a growing catalog of transportation and mobility-related data that empowers researchers and community planners to easily and securely share and preserve data that support projects and decision making.

Livewire, EEMS, energy-efficient mobility systems,↗

Compact, low power, high resolution ADC per pixel for large area pixel detectors

A compact ADC circuit can include one or more comparators, and a serial DAC (Digital-to-Analog) circuit that provides a signal to the comparator (or comparators). In addition, the ADC circuit can include a serial DAC redistribution sequencer that can provide a plurality of signals as input to the serial DAC circuit and is subject to a redistribution cycle and which receives as input a signal from a data multiplexer whose input connects electronically to an output of the comparator. The circuit can further include an ADC code register that provides an ADC output that connects electronically to the output of the comparator and the input to the data multiplexer. Shared logic circuitry for sharing common logic between pixels can be included, wherein the shared logic circuitry connects electronically to the data multiplexer and the ADC code register, wherein the shared logic circuitry promotes area and power savings for the pixel detector circuit.

Zimmerman, Tom↗

Compact, low power, high resolution ADC per pixel for large area pixel detectors

A compact ADC circuit can include one or more comparators, and a serial DAC (Digital-to-Analog) circuit that provides a signal to the comparator (or comparators). In addition, the ADC circuit can include a serial DAC redistribution sequencer that can provide a plurality of signals as input to the serial DAC circuit and is subject to a redistribution cycle and which receives as input a signal from a data multiplexer whose input connects electronically to an output of the comparator. The circuit can further include an ADC code register that provides an ADC output that connects electronically to the output of the comparator and the input to the data multiplexer. Shared logic circuitry for sharing common logic between pixels can be included, wherein the shared logic circuitry connects electronically to the data multiplexer and the ADC code register, wherein the shared logic circuitry promotes area and power savings for the pixel detector circuit.

Fahim, Farah↗

Privacy policy robustness to reverse engineering

Differential privacy policies allow one to preserve data privacy while sharing and analyzing data. However, these policies are susceptible to an array of attacks. In particular, often a portion of the data desired to be privacy protected is exposed online. Access to these pre-privacy protected data samples can then be used to reverse engineer the privacy policy. With knowledge of the generating privacy policy, an attacker can use machine learning to approximate the full set of originating data. Bayesian inference is one method for reverse engineering both model and model parameters. We present a methodology for evaluating and ranking privacy policy robustness to Bayesian inference-based reverse engineering, and demonstrated this method across data with a variety of temporal trends.

Kusne, Aaron Gilad↗

Rhizosphere Soil Biogeochemical Data and Photosynthetic Data of Vicia Faba in a Rhizobox

Here we share the data in column format via csv files for pH, redox, and dissolved oxygen collected at hourly resolution from microelectrodes. Dissolved organic carbon concentrations collected from TOC are also provided in a similar format but are composited samples from hourly microdialysis collection. This provided resolution of diel rhizosphere dynamics belowground. Plant physiological data was also collected at every 5 min for 24 hr cycles in order to capture diel dynamics aboveground. This data was used to parameterize the reaction transport model eSTOMP-ROOTS, which examines the rhizosphere biogeochemistry of a growing Vicia faba plant. The aim was to investigate plant activity and belowground biogeochemical processes, particularly their impact on mineral-organic associations in the rhizosphere. We combined in-situ rhizosphere microsensor and plant physiological measurements with a 3-D plant-soil reactive transport model to explore the behavior of dissolved organic carbon (DOC) in the rhizosphere. Over several days, microdialysis probes placed at the root-soil interface in live soil showed distinct daily patterns of DOC concentration in the pore water. Spikes in DOC concentrations during the day aligned with peaks in leaf-level photosynthesis, accompanied by decreasing redox potential and dissolved oxygen levels, and increasing pH in the rhizosphere. This new mechanistic modeling framework, which integrates aboveground plant physiological data with non-destructive, high-resolution monitoring of rhizosphere processes, offers significant potential for studying the factors that control carbon storage in soils.

54 ENVIRONMENTAL SCIENCES↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

MSD CoP Webinar: "Advances in MSD-LIVE to Support the MSD Community of Practice"

Context: This webinar was hosted by the MultiSector Dynamics Community of Practice (MSD CoP; https://multisectordynamics.org). Advances in MSD-LIVE to Support the MSD Community of Practice Presenters: Casey Burleyson and Zoe Guillen (Pacific Northwest National Laboratory) Abstract: The MultiSector Dynamics Living, Intuitive, Value-adding, Environment (MSD-LIVE; msdlive.org) is a cloud-based data management system and advanced computing platform that enables MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and workflows within the MSD Community of Practice. Recently, several high-profile datasets have attracted many new users to MSD-LIVE. This webinar has two goals: 1) To refamiliarize the MSD community and new users with the components of the platform (e.g., the data repository, model training notebooks, and data dashboards) and to highlight examples of how these components are advancing MSD science and 2) To demonstrate new features in v3 of the platform, released in late 2025. The main new feature in v3 is the ability to interactively explore data in MSD-LIVE without downloading it. MSD-LIVE users can now click a button in our data repository and launch a blank Jupyter notebook with access to the underlying data on AWS. Users can use the notebook to write analysis, visualization, or subsetting routines that process the data directly on the AWS cloud. We also added a GitHub integration feature that allows users to share analysis or visualization code they develop with the community of MSD-LIVE users. The webinar will wrap up with a look at what's coming next for MSD-LIVE in 2026. Moderator: Patrick M. Reed (MSD CoP Facilitation Team) This webinar was held on: May 12th, 2026 from 1-2 PM EST.

Open Science↗

Anonymization of Network Traces Data through Condensation-based Differential Privacy

Network traces are considered a primary source of information to researchers, who use them to investigate research problems such as identifying user behavior, analyzing network hierarchy, maintaining network security, classifying packet flows, and much more. However, most organizations are reluctant to share their data with a third party or the public due to privacy concerns. Therefore, data anonymization prior to sharing becomes a convenient solution to both organizations and researchers. Although several anonymization algorithms are available, few of them allow sufficient privacy (organization need), acceptable data utility (researcher need), and efficient data analysis at the same time. This article introduces a condensation-based differential privacy anonymization approach that achieves an improved tradeoff between privacy and utility compared to existing techniques and produces anonymized network trace data that can be shared publicly without lowering its utility value. Our solution also does not incur extra computation overhead for the data analyzer. A prototype system has been implemented, and experiments have shown that the proposed approach preserves privacy and allows data analysis without revealing the original data even when injection attacks are launched against it. When anonymized datasets are given as input to graph-based intrusion detection techniques, they yield almost identical intrusion detection rates as the original datasets with only a negligible impact.

97 MATHEMATICS AND COMPUTING↗

MOSIQS: Persistent Memory Object Storage With Metadata Indexing and Querying for Scientific Computing

Scientific applications often require high-bandwidth shared storage to perform joint simulations and collaborative data analytics. Shared memory pools provide a chance to satisfy such needs. Recently, a high-speed network such as Gen-Z utilizing persistent memory (PM) offers an opportunity to create a shared memory pool connected to compute nodes. However, there are several challenges to use scientific applications on the shared memory pool directly such as scalability, failure-atomicity, and lack of scientific metadata-based search and query. In this paper, we propose MOSIQS, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design MOSIQS based on the key idea that memory objects on PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. MOSIQS provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects to accelerate scientific computing. MOSIQS uses a lightweight persistent memory key-value store to manage the metadata of memory objects, which enables memory object sharing. To facilitate metadata search and query over millions of memory objects resident on memory pool, we introduce Group Split and Merge (GSM), a novel persistent index data structure designed primarily for scientific datasets. GSM splits and merges dynamically to minimize the query search space and maintains low query processing time while overcoming the index storage overhead. MOSIQS is implemented on top of PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. Experimental results show that MOSIQS gains a 100% write performance improvement and executes multi-attribute queries efficiently with 2.7× less index storage overhead offering significant potential to speed up scientific computing applications.

97 MATHEMATICS AND COMPUTING↗

Improving Discovery, Sharing, and Use of Water Data: Initial Findings and Suggested Future Work

Collaborative management of water resources requires a broad suite of “water data” that extends beyond basic information about water quantity and quality to other related topics such as water infrastructure, aquatic ecosystem health, socioeconomic factors, and power generation. Water data are disparate in nature because they are collected and provided by many entities, and in some cases, remain challenging to access and use. The U.S. Department of Energy’s Water Power Technologies Office initiated a project to characterize relevant categories of water data; describe the current state of accessing, using, and visualizing water data; and outline investigatory pathways for future efforts aimed at improving the discovery, sharing, and use of water data. Input on these topics was solicited from a small but diverse cross section of members of the water resources community. Fourteen broad categories of water data were identified: dams; ecology; flood control; hydroclimatology; hydrography; hydrology; hydropower; management landscape; migratory barriers; recreation and aesthetic importance; socioeconomic; water quality; water availability and use; and weather. Stakeholder perspectives on the accessibility and usability of water data indicate these aspects are affected by a complex set of technical and social factors. However, stakeholders generally agreed that better access to water data can provide a range of benefits to water management, and they stressed the need to generate broad support from water data users and producers. Two investigatory pathways were outlined that, taken together, provide a logical progression toward the goals of the project. The first pathway emphasizes further investigation to better define target audiences and data needs, identify opportunities for collaboration between related efforts, and conduct value demonstration activities to generate further support for improving discovery and access of water data. The second pathway focuses on creating a comprehensive vision for potential solutions that improve the discovery of water data. Several activities that align with the first pathway are suggested for the next phase of the project.

13 HYDRO ENERGY↗

msdlive-cli-distro

MSD-LIVE, the MultiSector Dynamics – Living, Intuitive, Value-adding, Environment, is a flexible and scalable data and code management system combined with a distributed computational platform that will enable MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and multi-model workflows within a robust Community of Practice. MSD-LIVE will facilitate a new open, collaborative, resource-rich, technology-facilitated, community-driven way of doing MSD research.

Lansing, Carina↗