Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “edge computer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Applications of Digital Twin System in a Smart City System with Multi-Energy

With the increasing number of electric vehicles (EV), the smart transportation system is becoming more closely related to the smart energy system in a smart city. However, because of the high complexity, dynamics, and large-scale of these two systems, it is challenging to study their operation problems systematically, especially in severe conditions, e.g., the 2021 winter Texas power traffic crisis. With artificial intelligence, federated learning, edge computing, and automatic control, a digital twin based smart city is proposed in this paper and it focuses on two applications: smart transportation and smart energy grid. Both systems are presented in a structured design manner with several components in the digital twin system, which also contains many real-implementations of different scenarios. The proposed digital twin of a smart city provides a potential clue for the problems mentioned above.

artificial intelligence↗

Perspectives on AI Architectures and Codesign for Earth System Predictability

Abstract Recently, the U.S. Department of Energy (DOE), Office of Science, Biological and Environmental Research (BER), and Advanced Scientific Computing Research (ASCR) programs organized and held the Artificial Intelligence for Earth System Predictability (AI4ESP) workshop series. From this workshop, a critical conclusion that the DOE BER and ASCR community came to is the requirement to develop a new paradigm for Earth system predictability focused on enabling artificial intelligence (AI) across the field, laboratory, modeling, and analysis activities, called model experimentation (ModEx). BER’s ModEx is an iterative approach that enables process models to generate hypotheses. The developed hypotheses inform field and laboratory efforts to collect measurement and observation data, which are subsequently used to parameterize, drive, and test model (e.g., process based) predictions. A total of 17 technical sessions were held in this AI4ESP workshop series. This paper discusses the topic of the AI Architectures and Codesign session and associated outcomes. The AI Architectures and Codesign session included two invited talks, two plenary discussion panels, and three breakout rooms that covered specific topics, including 1) DOE high-performance computing (HPC) systems, 2) cloud HPC systems, and 3) edge computing and Internet of Things (IoT). We also provide forward-looking ideas and perspectives on potential research in this codesign area that can be achieved by synergies with the other 16 session topics. These ideas include topics such as 1) reimagining codesign, 2) data acquisition to distribution, 3) heterogeneous HPC solutions for integration of AI/ML and other data analytics like uncertainty quantification with Earth system modeling and simulation, and 4) AI-enabled sensor integration into Earth system measurements and observations. Such perspectives are a distinguishing aspect of this paper. Significance Statement This study aims to provide perspectives on AI architectures and codesign approaches for Earth system predictability. Such visionary perspectives are essential because AI-enabled model-data integration has shown promise in improving predictions associated with climate change, perturbations, and extreme events. Our forward-looking ideas guide what is next in codesign to enhance Earth system models, observations, and theory using state-of-the-art and futuristic computational infrastructure.

54 ENVIRONMENTAL SCIENCES↗

Perspectives on AI Architectures and Co-design for Earth System Predictability

Recently, the U.S. Department of Energy (DOE), Office of Science, Biological and Environmental Research (BER), and Advanced Scientific Computing Research (ASCR) programs organized and held the Artificial Intelligence for Earth System Predictability (AI4ESP) workshop series. From this workshop, a critical conclusion that the DOE BER and ASCR community came to is the requirement to develop a new paradigm for Earth system predictability focused on enabling artificial intelligence (AI) across the field, laboratory, modeling, and analysis activities, called model experimentation (ModEx). BER’s ModEx is an iterative approach that enables process models to generate hypotheses. The developed hypotheses inform field and laboratory efforts to collect measurement and observation data, which are subsequently used to parameterize, drive, and test model (e.g., process based) predictions. A total of 17 technical sessions were held in this AI4ESP workshop series. This paper discusses the topic of the AI Architectures and Codesign session and associated outcomes. The AI Architectures and Codesign session included two invited talks, two plenary discussion panels, and three breakout rooms that covered specific topics, including 1) DOE high-performance computing (HPC) systems, 2) cloud HPC systems, and 3) edge computing and Internet of Things (IoT). We also provide forward-looking ideas and perspectives on potential research in this codesign area that can be achieved by synergies with the other 16 session topics. These ideas include topics such as 1) reimagining codesign, 2) data acquisition to distribution, 3) heterogeneous HPC solutions for integration of AI/ML and other data analytics like uncertainty quantification with Earth system modeling and simulation, and 4) AI-enabled sensor integration into Earth system measurements and observations. Such perspectives are a distinguishing aspect of this paper.

58 GEOSCIENCES↗

Discrete Roughness Effects on Shuttle Orbiter at Mach 6

Discrete roughness boundary layer transition results on a Shuttle Orbiter model in the NASA Langley Research Center 20-Inch Mach 6 Air Tunnel have been reanalyzed with new boundary layer calculations to provide consistency for comparison to other published results. The experimental results were previously obtained utilizing the phosphor thermography system to monitor the status of the boundary layer via global heat transfer images of the Orbiter windward surface. The size and location of discrete roughness elements were systematically varied along the centerline of the 0.0075-scale model at an angle of attack of 40 deg and the boundary layer response recorded. Various correlative approaches were attempted, with the roughness transition correlations based on edge properties providing the most reliable results. When a consistent computational method is used to compute edge conditions, transition datasets for different configurations at several angles of attack have been shown to collapse to a well-behaved correlation.

Berry, Scott A.↗

When in-memory computing meets spiking neural networks—A perspective on device-circuit-system-and-algorithm co-design

This review explores the intersection of bio-plausible artificial intelligence in the form of spiking neural networks (SNNs) with the analog in-memory computing (IMC) domain, highlighting their collective potential for low-power edge computing environments. Through detailed investigation at the device, circuit, and system levels, we highlight the pivotal synergies between SNNs and IMC architectures. Additionally, we emphasize the critical need for comprehensive system-level analyses, considering the inter-dependencies among algorithms, devices, circuit, and system parameters, crucial for optimal performance. An in-depth analysis leads to the identification of key system-level bottlenecks arising from device limitations, which can be addressed using SNN-specific algorithm–hardware co-design techniques. This review underscores the imperative for holistic device to system design-space co-exploration, highlighting the critical aspects of hardware and algorithm research endeavors for low-power neuromorphic solutions.

Physics↗

A Three-Dimensional Solution of Flows over Wings with Leading-Edge Vortex Separation. Part 1: Engineering Document

A method of predicting forces, moments, and detailed surface pressures on thin, sharp-edged wings with leading-edge vortex separation in incompressible flow is presented. The method employs an inviscid flow model in which the wing and the rolled-up vortex sheets are represented by piecewise, continuous quadratic doublet sheet distributions. The Kutta condition is imposed on all wing edges. Computed results are compared with experimental data and with the predictions of the leading-edge suction analogy for a selected number of wing planforms over a wide range of angle of attack. These comparisons show the method to be very promising, capable of producing not only force predictions, but also accurate predictions of detailed surface pressure distributions, loads, and moments.

Brune, G. W.↗

GAHLS: an optimized graph analytics based high level synthesis framework

The urgent need for low latency, high-compute and low power on-board intelligence in autonomous systems, cyber-physical systems, robotics, edge computing, evolvable computing, and complex data science calls for determining the optimal amount and type of specialized hardware together with reconfigurability capabilities. With these goals in mind, we propose a novel comprehensive graph analytics based high level synthesis (GAHLS) framework that efficiently analyzes complex high level programs through a combined compiler-based approach and graph theoretic optimization and synthesizes them into message passing domain-specific accelerators. This GAHLS framework first constructs a compiler-assisted dependency graph (CaDG) from low level virtual machine (LLVM) intermediate representation (IR) of high level programs and converts it into a hardware friendly description representation. Next, the GAHLS framework performs a memory design space exploration while account for the identified computational properties from the CaDG and optimizing the system performance for higher bandwidth. The GAHLS framework also performs a robust optimization to identify the CaDG subgraphs with similar computational structures and aggregate them into intelligent processing clusters in order to optimize the usage of underlying hardware resources. Finally, the GAHLS framework synthesizes this compressed specialized CaDG into processing elements while optimizing the system performance and area metrics. Evaluations of the GAHLS framework on several real-life applications (e.g., deep learning, brain machine interfaces) demonstrate that it provides 14.27× performance improvements compared to state-of-the-art approaches such as LegUp 6.2.

97 MATHEMATICS AND COMPUTING↗

Synthetic-domain computing and neural networks using lithium niobate integrated nonlinear phononics

Analogue computing uses the physical behaviours of devices to provide energy-efficient arithmetic operations. However, scaling up analogue computing platforms by simply increasing the number of devices leads to challenges such as device-to-device variation. Here, in this study, we report scalable analogue computing and neural networks in the synthetic frequency domain using an integrated nonlinear phononic platform on lithium niobate. This synthetic-domain computing is robust to device variations, as vectors and matrices are concurrently encoded at different frequencies within a single device, achieving a high throughput per area. Leveraging inherent nonlinearities, our device-aware neural network can perform a four-class classification task with an accuracy of 98.2%. The nonlinear phononic computing hardware also maintains consistent performance over a wide operational temperature range (characterized up to 192 °C). Our synthetic-domain computing combines single-device parallelism, inherent nonlinearity and environmental stability, and could be of use in edge computing applications in which power efficiency and environmental resilience are crucial.

Ji, Jun [Virginia Polytechnic Inst. and State Univ↗

Uncovering acoustic signatures of pore formation in laser powder bed fusion

Abstract We present a machine learning workflow to discover signatures in acoustic measurements that can be utilized to create a low-dimensional model to accurately predict the location of keyhole pores formed during additive manufacturing processes. Acoustic measurements were sampled at 100 kHz during single-layer laser powder bed fusion (LPBF) experiments, and spatio-temporal registration of pore locations was obtained from post-build radiography. Power spectral density (PSD) estimates of the acoustic data were then decomposed using non-negative matrix factorization with custom $$\varvec{k}$$ k -means clustering (NMF $$\varvec{k}$$ k ) to learn the underlying spectral patterns associated with pore formation. NMF $$\varvec{k}$$ k returned a library of basis signals and matching coefficients to blindly construct a feature space based on the PSD estimates in an optimized fashion. Moreover, the NMF $$\varvec{k}$$ k decomposition led to the development of computationally inexpensive machine learning models which are capable of quickly and accurately identifying pore formation with classification accuracy of supervised and unsupervised label learning greater than 95% and 90%, respectively. The intrinsic data compression of NMF k , the relatively light computational cost of the machine learning workflow, and the high classification accuracy makes the proposed workflow an attractive candidate for edge computing toward in-situ keyhole pore prediction in LPBF.

36 MATERIALS SCIENCE↗

Digital Modeling on Large Kernel Metamaterial Neural Network

Deep neural networks (DNNs) utilized recently are physically deployed with computational units (e.g., CPUs and GPUs). Such a design might lead to a heavy computational burden, significant latency, and intensive power consumption, which are critical limitations in applications such as Internet of Things (IoT), edge computing, and usage of drones. Recent advances in optical computational units (e.g., metamaterial) have shed light on energy-free and light-speed neural networks. However, the digital design of the metamaterial neural network (MNN) is fundamentally limited by its physical limitations, such as precision, noise, and bandwidth during fabrication. Moreover, the unique advantages of MNN’s (e.g., light-speed computation) are not fully explored via standard 3×3 convolution kernels. In this paper, we propose a novel large kernel metamaterial neural network (LMNN) that maximizes the digital capacity of the state-of-the-art (SOTA) MNN with model re-parametrization and network compression, while also considering the optical limitation explicitly. The new digital learning scheme can maximize the learning capacity of MNN while modeling the physical restrictions of meta-optics. With the proposed LMNN, the computation cost of the convolutional front-end can be offloaded to fabricated optical hardware. The experimental results on two publicly available datasets demonstrate that the optimized hybrid design improved classification accuracy while reducing computational latency. In conclusion, the development of the proposed LMNN is a promising step towards the ultimate goal of energy-free and light-speed AI.

97 MATHEMATICS AND COMPUTING↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Leveraging Artificial Intelligence to Expedite Antibody Design and Enhance Antibody–Antigen Interactions

This perspective sheds light on the transformative impact of recent computational advancements in the field of protein therapeutics, with a particular focus on the design and development of antibodies. Cutting-edge computational methods have revolutionized our understanding of protein–protein interactions (PPIs), enhancing the efficacy of protein therapeutics in preclinical and clinical settings. Central to these advancements is the application of machine learning and deep learning, which offers unprecedented insights into the intricate mechanisms of PPIs and facilitates precise control over protein functions. Despite these advancements, the complex structural nuances of antibodies pose ongoing challenges in their design and optimization. Our review provides a comprehensive exploration of the latest deep learning approaches, including language models and diffusion techniques, and their role in surmounting these challenges. We also present a critical analysis of these methods, offering insights to drive further progress in this rapidly evolving field. The paper includes practical recommendations for the application of these computational techniques, supplemented with independent benchmark studies. These studies focus on key performance metrics such as accuracy and the ease of program execution, providing a valuable resource for researchers engaged in antibody design and development. Through this detailed perspective, we aim to contribute to the advancement of antibody design, equipping researchers with the tools and knowledge to navigate the complexities of this field.

59 BASIC BIOLOGICAL SCIENCES↗

Proton Tunable Analog Transistor for Low Power Computing

This project was broadly motivated by the need for new hardware that can process information such as images and sounds right at the point of where the information is sensed (e.g. edge computing). The project was further motivated by recent discoveries by group demonstrating that while certain organic polymer blends can be used to fabricate elements of such hardware, the need to mix ionic and electronic conducting phases imposed limits on performance, dimensional scalability and the degree of fundamental understanding of how such devices operated. As an alternative to blended polymers containing distinct ionic and electronic conducting phases, in this LDRD project we have discovered that a family of mixed valence coordination compounds called Prussian blue analogue (PBAs), with an open framework structure and ability to conduct both ionic and electronic charge, can be used for inkjet-printed flexible artificial synapses that reversibly switch conductance by more than four orders of magnitude based on electrochemically tunable oxidation state. Retention of programmed states is improved by nearly two orders of magnitude compared to the extensively studied organic polymers, thus enabling in-memory compute and avoiding energy costly off-chip access during training. We demonstrate dopamine detection using PBA synapses and biocompatibility with living neurons, evoking prospective application for brain - computer interfacing. By application of electron transfer theory to in-situ spectroscopic probing of intervalence charge transfer, we elucidate a switching mechanism whereby the degree of mixed valency between N-coordinated Ru sites controls the carrier concentration and mobility, as supported by density functional theory (DFT) .

97 MATHEMATICS AND COMPUTING↗

Measurement and applications: Exploring the challenges and opportunities of hierarchical federated learning in sensor applications

Sensor applications have become ubiquitous in modern society as the digital age continues to advance. AI-based techniques (e.g., machine learning) are effective at extracting actionable information from large amounts of data. An example would be an automated water irrigation system that uses AI-based techniques on soil quality data to decide how to best distribute water. However, these AI-based techniques are costly in terms of hardware resources, and Internet-of-Things (IoT) sensors are resource-constrained with respect to processing power, energy, and storage capacity. These limitations can compromise the security, performance, and reliability of sensor-driven applications. To address these concerns, cloud computing services can be used by sensor applications for data storage and processing. Unfortunately, cloud-based sensor applications that require real-time processing, such as medical applications (e.g., fall detection and stroke prediction), are vulnerable to issues such as network latency due to the sparse and unreliable networks between the sensor nodes and the cloud server [1]. As users approach the edge of the communications network, latency issues become more severe and frequent. A promising alternative is edge computing, which provides cloud-like capabilities at the edge of the network by pushing storage and processing capabilities from centralized nodes to edge devices that are closer to where the data are gathered, resulting in reduced network delays [2], [3].

Po-Leen Ooi, Melanie↗

Analyzing Data Privacy for Edge Systems

Internet-of-Things (IoT)-based streaming applications are all around us. Currently, we are transitioning from IoT processing being performed on the cloud to the edge. While moving to the edge provides significant networking efficiency benefits, IoT edge computing creates significant data privacy concerns. We propose a methodology that can successfully privacy protect the continual data streams generated by sensors on the edge device. We implement local differential privacy on streaming data and incorporate Bayesian inference and Gaussian process to evaluate the privacy policy. We demonstrate our methodology on a real-world smart meter testbed and identify the optimal privacy protection settings.

Kotevska, Olivera↗

Reimagining Codesign for Advanced Scientific Computing: Report for the ASCR Workshop on Reimagining Codesign

In March 2021, the U.S. Department of Energy’s Advanced Scientific Computing Research program convened the Workshop on Reimagining Codesign. The workshop, also known as ReCoDe, was organized around discussions on eight topic areas: (1) codesign for traditional high-performance computing workloads; (2) codesign of memory/storage systems; (3) codesign of machine learning, neuromorphic, quantum, and other non-von Neumann accelerators; (4) codesign for edge computing and processing at experimental instruments; (5) codesign for security and privacy; (6) hardware design tools and open-source hardware for high-productivity codesign; (7) tools, software stack, and programming languages for high-productivity codesign; and (8) quantitative tools and data collection for modeling and simulation for codesign. The panels identified four Priority Research Directions from these deliberations: (1) breakthrough computing capabilities with targeted heterogeneity and rapid design; (2) software and applications that embrace radical architecture diversity; (3) engineered security and integrity, from transistors to applications; and (4) design with data-rich processes.

97 MATHEMATICS AND COMPUTING↗

Scientific computing

Our research generates a vast amount of data that scientists must gather, store, analyze and interpret. These activities require expertise and the right tools. Fermilab is home to one of the largest tape robotic systems available today, providing more than 800 petabytes of storage capability. The lab is growing the tape facilities to meet expected demands of over 1,000 petabytes of additional data per year by the end of the decade. Fermilab also has the technology and computing power to efficiently pr ocess these data to facilitate scientific discoveries. Our computation al experts work closely with experimenters and the scientific commu - nity at large to provide and facilitate the use of the cutting-edge computing tools necessary for these discoveries.

Fermilab, Fermilab↗

Building an Integrated Ecosystem of Computational and Observational Facilities to Accelerate Scientific Discovery

Future scientific discoveries will rely on flexible ecosystems that incorporate modern scientific instruments, high performance computing resources, parallel distributed data storage, and performant networks across multiple, independent facilities. In addition to connecting physical resources, such an ecosystem presents many challenges in logistics and accessibility, especially in orchestrating computations and experiments that span across leadership computing systems and experimental instruments. Past efforts have typically been application-specific or limited to interfaces for computing resources. This paper proposes a general framework for integrating computation resources and instrument operations, addressing challenges in code development/execution, data staging and collection, software stack, control mechanisms, resource authorization and governance, and hardware integration. We also describe a demonstration use case wherein a Bayesian optimization algorithm running on an edge computing resource guides a scanning probe microscope to autonomously and intelligently characterize a material sample. This science edge ecosystem framework will provide a blueprint for federating multi-institutional, disparate resources and orchestrating scientific workflows across them to enable next-generation discoveries.

Somnath, Suhas↗