Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “secure data transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data transmission by quantum matter wave modulation

Abstract Classical communication schemes exploiting wave modulation are the basis of our information era. Quantum information techniques with photons enable future secure data transfer in the dawn of decoding quantum computers. Here we demonstrate that also matter waves can be applied for secure data transfer. Our technique allows the transmission of a message by a quantum modulation of coherent electrons in a biprism interferometer. The data is encoded in the superposition state by a Wien filter introducing a longitudinal shift between separated matter wave packets. The transmission receiver is a delay line detector performing a dynamic contrast analysis of the fringe pattern. Our method relies on the Aharonov–Bohm effect but does not shift the phase. It is demonstrated that an eavesdropping attack will terminate the data transfer by disturbing the quantum state and introducing decoherence. Furthermore, we discuss the security limitations of the scheme due to the multi-particle aspect and propose the implementation of a key distribution protocol that can prevent active eavesdropping.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

3D printed graphene-based self-powered strain sensors for smart tires in autonomous vehicles

The transition of autonomous vehicles into fleets requires an advanced control system design that relies on continuous feedback from the tires. Smart tires enable continuous monitoring of dynamic parameters by combining strain sensing with traditional tire functions. Here, we provide breakthrough in this direction by demonstrating tire-integrated system that combines direct mask-less 3D printed strain gauges, flexible piezoelectric energy harvester for powering the sensors and secure wireless data transfer electronics, and machine learning for predictive data analysis. Ink of graphene based material was designed to directly print strain sensor for measuring tire-road interactions under varying driving speeds, normal load, and tire pressure. A secure wireless data transfer hardware powered by a piezoelectric patch is implemented to demonstrate self-powered sensing and wireless communication capability. Combined, this study significantly advances the design and fabrication of cost-effective smart tires by demonstrating practical self-powered wireless strain sensing capability.

33 ADVANCED PROPULSION SYSTEMS↗

Regression Analysis with the Directed Infusion of Data

Integrating artificial intelligence and machine learning tools into industry necessitates large-scale collaborative efforts that ensure the robust and accurate execution of downstream analytics such as time series prediction, uncertainty quantification, grid optimization, and condition monitoring. However, concerns related to data privacy pervade the nuclear industry due to the proprietary nature of its data and the possibility of data leakage. Legacy techniques such as encryption often require the explicit transmission of data to trustworthy parties, thereby inviting data leakage concerns. The ideal collaboration scenario avoids the explicit dissemination of data/code while maintaining experimental fidelity, which is currently accomplished using various techniques such as trusted execution environments, homomorphic encryption, differential privacy, and multimatrix masking. These techniques, however, often necessitate a trade-off between trust, efficiency, and utility. This article extends a previously proposed technique called the directed infusion of data (DIOD) that ensures data privacy, allows for scalable obfuscation, and combats the risk of data leakage without compromising utility. The experiments discussed in this article examine a regression-type scenario using DIOD with the goal of preserving the inferential link between two variables. Using the point-kinetics equations, regression experiments compare the performance of a model trained using the original data to that of a model trained using the obfuscated data, which produced identical results. Our claim is further strengthened by an information theoretic proof and experiment, which showed that the inferential content between variables remains the same after obfuscation, thereby avoiding the required communication of the proprietary data.

47 - OTHER INSTRUMENTATION↗

Reconfigurable Network Slicing Orchestration in Network Function Virtualization Compatible Operational Technology Environment

The ongoing transition to Industry 4.0, which is characterized by increased inter-connectivity of cyber-physical systems, requires having time-sensitive, high throughput, and secure transfer of critical data in industrial sites. In this context, network slicing emerges as a critical tool to ensure timely data delivery by provisioning the network resources to cater to specific applications’ requirements and mitigating potential cyber attacks. To address these challenges, this paper aims to tackle two key questions essential for the successful implementation of network slicing in industrial environments. First, it investigates architectural considerations for developing a network infrastructure capable of supporting network slicing functionalities effectively. The proposed approach significantly improves deployment efficiency over traditional manual configurations. Second, it delves into the automated orchestration process, elucidating the steps and components involved in transitioning from a static network management approach to dynamically leverage network function virtualization schemes for creating network slices in ad-hoc manner. The system demonstrates high throughput suitable for production-level solutions and maintains exceptionally low latency, making it ideal for ultra-reliable low-latency communications. Even with increased network demands, the system remains stable, with effective Quality of Service (QoS) management, ensuring reliable performance under varying conditions. The proposed architecture outlines the necessary components, services, and communication protocols required for a production-level orchestrator for network segmentation in SCADA environments.

Rodiles Delgado, Brian G.↗

Joint Genome Institute Analysis Workflow Service (JAWS) v2.7

The U.S. Department of Energy Joint Genome Institute (JGI) has developed the JGI Analysis Workflow Service (JAWS) as a distributed framework to run computational workflows across diverse high-performance computing (HPC) and cloud environments. JAWS enhances the reusability, scalability, and robustness of scientific workflows by orchestrating data movement, code execution, and results retrieval across multiple DOE facilities. At its core, JAWS integrates the Cromwell workflow engine to run workflows expressed in the Workflow Description Language (WDL), ensuring portability and interoperability. To provide consistent runtime environments, JAWS employs container technologies such as Shifter, Apptainer, and Docker. Workflow tasks are managed via HTCondor on HPC backends, while Globus ensures secure and efficient data transfer between sites. JAWS is deployed as a multi-site workflow manager across national laboratory computing facilities, with dedicated instances supporting community projects such as the National Microbiome Data Collaborative (NMDC) and KBase. This distributed, service-oriented architecture enables users to "write once, run anywhere," providing scalable, production-quality workflow execution.

Kirton, Edward↗

One-way transfer device with secure reverse channel

A data diode provides a flexible device for collecting data from a data source and transmitting the data to a data destination using one-way data transmission across a main channel. On-board processing elements allow the data diode to identify automatically the type of connectivity provided to the data diode and configure the data diode to handle the identified type of connectivity. Either or both of the inbound and outbound side of the data diode may comprise one or both of wired and wireless communication interfaces. A secure reverse channel, separate from the main channel, allows carefully predetermined communications from the data destination to the data source.

Lee, Sang Cheon↗

One-way transfer device with secure reverse channel

A data diode provides a flexible device for collecting data from a data source and transmitting the data to a data destination using one-way data transmission across a main channel. On-board processing elements allow the data diode to identify automatically the type of connectivity provided to the data diode and configure the data diode to handle the identified type of connectivity. Either or both of the inbound and outbound side of the data diode may comprise one or both of wired and wireless communication interfaces. A secure reverse channel, separate from the main channel, allows carefully predetermined communications from the data destination to the data source.

Lee, Sang Cheon↗

Data transfer for STAR grid jobs

The Solenoidal Tracker at RHIC (STAR) is a multipurpose experiment at the Relativistic Heavy Ion Collider (RHIC) with the primary goal to study the formation and properties of the quark-gluon plasma. STAR is an international collaboration of member institutions and laboratories from around the world. Yearly data-taking period produces PBytes of raw data collected by the experiment. STAR primarily uses its dedicated facility at BNL to process this data, but has routinely leveraged distributed systems, both high throughput (HTC) and high performance (HPC) computing clusters, to significantly augment the processing capacity available to the experiment. The ability to automate the efficient transfer of large data sets on reliable, scalable, and secure infrastructure is critical for any large-scale distributed processing campaign. For more than a decade, STAR computing has relied upon GridFTP with its x509-based authentication to build such data transfer systems and integrate them into its larger production workflow. The end of support by the community for both GridFTP and the x509 standard requires STAR to investigate other approaches to meet its distributed processing needs. In this study we investigate two multi-purpose data distribution systems, Globus.org and XRootD, as alternatives to GridFTP. We compare both their performance and the ease by which each service is integrated into the type of secure and automated data transfer systems STAR has previously built using GridFTP. The presented approach and study may be applicable to other distributed data processing use cases beyond STAR.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Automated, reliable, and efficient continental-scale replication of 7.3 petabytes of computational simulation data: A case study

We report on our experiences replicating 7.3 petabytes (PB) of Earth System Grid Federation (ESGF) computational simulation data from Lawrence Livermore National Laboratory (LLNL) in California to Argonne National Laboratory (ANL) in Illinois and Oak Ridge National Laboratory (ORNL) in Tennessee—a task motivated by a need for increased reliability, capacity, and performance. This task presented significant challenges: the need to move 29 million files twice under time pressure from aging storage hardware; a source file system bottleneck limiting throughput to 1.5 GB/s; frequent site maintenance windows; and the need for complete reliability at scale. We addressed these challenges using a simple replication tool that invoked Globus to transfer large bundles of files while tracking progress in a database, dynamically rerouting transfers to work around maintenance periods and file system limitations. Under the covers, Globus organized transfers to make efficient use of the high-speed Energy Sciences network (ESnet) and the data transfer nodes deployed at participating sites, and also addressed security, integrity checking, and recovery from a variety of transient failures. This success demonstrates the considerable benefits that can accrue from the adoption of performant data replication infrastructure. The replication tool is available at https://github.com/esgf2-us/data-replication-tools.

Globus↗

Federated learning for 2D synchrotron x-ray diffractometry: a cross-institutional approach for phase quantification of Ti–6Al–4V alloy

High-energy Two dimensional (2D) synchrotron x-ray diffractometry provides important insights into the atomistic structure and phase evolution of materials, yet traditional analysis methods remain complex, knowledge-intensive, and computationally demanding. Deep-learning models offer a powerful alternative for automating their analysis. Institutions that hold these datasets may be unwilling to share their data due to privacy and security policies, as well as the challenges associated with large-scale data transfer. As a result, models trained on local datasets often perform well only on their own data but exhibit bias and poor generalization across different instruments or facilities. To overcome these limitations, we explore federated learning (FL) for 2D synchrotron diffractograms, enabling collaborative model training without exchanging raw data. In this study, 2D synchrotron diffractograms of Ti–6Al–4V alloy collected from two independent facilities are used to train convolutional neural networks for predicting the β-phase volume fraction. Experimental results show that federated global models significantly outperform locally trained models in terms of generalization and achieve accuracy comparable to centralized trained models. These findings demonstrate the potential of FL to enable secure, cross-institutional collaboration and enhance the scalability of deep-learning-based materials characterization.

36 MATERIALS SCIENCE↗

Privacy-Preserving Knowledge Transfer with Bootstrap Aggregation of Teacher Ensembles

There is a need to transfer knowledge among institutions and organizations to save effort in annotation and labeling or in enhancing task performance. However, knowledge transfer is difficult because of restrictions that are in place to ensure data security and privacy. Institutions are not allowed to exchange data or perform any activity that may expose personal information. With the leverage of a differential privacy algorithm in a high-performance computing environment, we propose a new training protocol, Bootstrap Aggregation of Teacher Ensembles (BATE), which is applicable to various types of machine learning models. The BATE algorithm is based on and provides enhancements to the PATE algorithm, maintaining competitive task performance scores on complex datasets with underrepresented class labels.We conducted a proof-of-the-concept study of the information extraction from cancer pathology report data from four cancer registries and performed comparisons between four scenarios: no collaboration, no privacy-preserving collaboration, the PATE algorithm, and the proposed BATE algorithm. The results showed that the BATE algorithm maintained competitive macro-averaged F1 scores, demonstrating that the suggested algorithm is an effective yet privacy-preserving method for machine learning and deep learning solutions.

Yoon, Hong-Jun↗

Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments

There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.

97 MATHEMATICS AND COMPUTING↗

Fed-DeepONet: Stochastic Gradient-Based Federated Training of Deep Operator Networks

The Deep Operator Network (DeepONet) framework is a different class of neural network architecture that one trains to learn nonlinear operators, i.e., mappings between infinite-dimensional spaces. Traditionally, DeepONets are trained using a centralized strategy that requires transferring the training data to a centralized location. Such a strategy, however, limits our ability to secure data privacy or use high-performance distributed/parallel computing platforms. To alleviate such limitations, in this paper, we study the federated training of DeepONets for the first time. That is, we develop a framework, which we refer to as Fed-DeepONet, that allows multiple clients to train DeepONets collaboratively under the coordination of a centralized server. To achieve Fed-DeepONets, we propose an efficient stochastic gradient-based algorithm that enables the distributed optimization of the DeepONet parameters by averaging first-order estimates of the DeepONet loss gradient. Then, to accelerate the training convergence of Fed-DeepONets, we propose a moment-enhanced (i.e., adaptive) stochastic gradient-based strategy. Finally, we verify the performance of Fed-DeepONet by learning, for different configurations of the number of clients and fractions of available clients, (i) the solution operator of a gravity pendulum and (ii) the dynamic response of a parametric library of pendulums.

Moya, Christian↗

Development of A Hardware-In-the-Loop (HIL) Testbed for Cyber-Physical Security in Smart Buildings

As smart buildings move towards open communication technologies, providing access to the Building Automation System (BAS) through the building's intranet, or even remotely through the Internet, has become a common practice. However, BAS was historically developed as a closed environment and designed with limited cyber-security considerations. Thus, smart buildings are vulnerable to cyber-attacks with the increased accessibility. This study introduces the development and capability of a Hardware-in-the-Loop (HIT) testbed for testing and evaluating the cyber-physical security of typical BASs in smart buildings. The testbed consists of three subsystems: (1) a real-time HIL emulator simulating the behavior of a virtual building as well as the Heating, Ventilation, and Air Conditioning (HVAC) equipment via a dynamic simulation in Modelica; (2) a set of real HVAC controllers monitoring the virtual building operation and providing local control signals to control HVAC equipment in the HIL emulator; and (3) a BAS server along with a web-based service for users to fully access the schedule, setpoints, trends, alarms, and other control functions of the HVAC controllers remotely through the BACnet network. The server generates rule-based setpoints to local HVAC controllers. Based on these three subsystems, the HIL testbed supports attack/fault-free and attack/fault-injection experiments at various levels of the building system. The resulting test data can be used to inform the building community and support the cyber-physical security technology transfer to the building industry.

Li, Guowen↗

ESnet Secure Copy (EScp) v0.6

EScp is a high speed transfer tool with a similar command line syntax to scp. Unlike SCP it is designed to transfer files at high speed, thus far we have been able to show 100gbit/s transfers, although I expect that the throughput should scale in proportion to the network interface, i.e. I expect 400gbit/s performance on our 400gbit/s test bed. EScp achieves good performance through an innovative design (multithreaded, zero copy transfers), along with pluggable filters and I/O engines. As an example, you can switch from POSIX i/O to UIO by checking a different engine. It also natively supports encryption, and cheksums for file verification and transport security. AAA is through standard SSH (same as SCP). By taking advantage of filters, EScp supports transferring unstructured data and/or I/O to non-posix data sources. Examples include streaming data (i.e. from equipment), transferring data to the cloud, and/or supporting non-posix file systems (like HPSS).

Shiflett, Charles↗

Unique contributions of chlorophyll and nitrogen to predict crop photosynthetic capacity from leaf spectroscopy

The photosynthetic capacity or the CO 2 -saturated photosynthetic rate (Vmax), chlorophyll, and nitrogen are closely linked leaf traits that determine C 4 crop photosynthesis and yield. Accurate, timely, rapid, and non-destructive approaches to predict leaf photosynthetic traits from hyperspectral reflectance are urgently needed for high-throughput crop monitoring to ensure food and bioenergy security. Therefore, this study thoroughly evaluated the state-of-the-art physically based radiative transfer models (RTMs), data-driven partial least squares regression (PLSR), and generalized PLSR (gPLSR) models to estimate leaf traits from leaf-clip hyperspectral reflectance, which was collected from maize (Zea mays L.) bioenergy plots with diverse genotypes, growth stages, treatments with nitrogen fertilizers, and ozone stresses in three growing seasons. The results show that leaf RTMs considering bidirectional effects can give accurate estimates of chlorophyll content (Pearson correlation r=0.95), while gPLSR enabled retrieval of leaf nitrogen concentration (r=0.85). Using PLSR with field measurements for training, the cross-validation indicates that V max can be well predicted from spectra (r=0.81). Here, the integration of chlorophyll content (strongly related to visible spectra) and nitrogen concentration (linked to shortwave infrared signals) can provide better predictions of V max (r=0.71) than only using either chlorophyll or nitrogen individually. This study highlights that leaf chlorophyll content and nitrogen concentration have key and unique contributions to V max prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Secure Federated Learning Across Heterogeneous Cloud and High-Performance Computing Resources: A Case Study on Federated Fine-Tuning of LLaMA 2

Federated learning enables multiple data owners to collaboratively train robust machine learning models without transferring large or sensitive local datasets by only sharing the parameters of the locally trained models. Here, in this article, we elaborate on the design of our Advanced Privacy-Preserving Federated Learning (APPFL) framework, which streamlines end-to-end secure and reliable federated learning experiments across cloud computing facilities and high-performance computing resources by leveraging Globus Compute, a distributed function as a service platform, and Amazon Web Services. We further demonstrate the use case of APPFL in fine-tuning an LLaMA 2 7B model using several cloud resources and supercomputers.

97 MATHEMATICS AND COMPUTING↗