Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Edge Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Light-powered end-to-end neutron detection and imaging with an edge-deployed optical AI chip

Neutron detection is widely used in many applications including nuclear physics, nuclear energy, nuclear technologies and nuclear safeguards. Developing an end-to-end neutron detection and imaging workflow paves way towards fully automated processes for many applications. We implemented an automated workflow for neutron detection experiments which use a solid state image sensor to capture neutron hits as a digital image. We deploy the workflow to an edge-based optical neural network (ONN) to increase the radiation-hardness and lifetime of neutron detection instruments. We present a two-stage neural network framework for detection of neutrons at sub-pixel resolution. The first stage uses a region proposal network to efficiently detect and extract neutron hits from the input camera image. The second stage feeds the extracted hits into a fully connected neural network to predict the sub-pixel hit position. The performance of the two-stage framework is evaluated using the edge-based ONN. The results show that we can achieve above 96% neutron detection accuracy as well as sub-pixel and sub-micron position resolution, while enjoying the advantages of the ONN hardware including radiation-hardness, low energy consumption and high computing speed for integrated edge camera and hardware deployment, when compared with electronic counterparts.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

Computationally efficient method for determining limiting velocities of edge dislocations in anisotropic crystals

The continuum-limit theory of dislocations in crystals predicts divergences in the elastic energy at crystal-geometry dependent limiting velocities vL, which separate subsonic, transsonic, and supersonic dislocation glide regimes and are therefore import for material strength models at high strain rates. Although it is known how to calculate those limiting velocities, there is one special case - edge dislocations with reflection symmetry, but non-vanishing elastic constants c16 or c26 - where previous methods have been notoriously slow. In this letter, we address this deficiency by deriving a computationally efficient method for determining the limiting velocities of edge dislocations with reflection symmetry which is two orders of magnitude faster than the previous method.

36 MATERIALS SCIENCE↗

A Flang Plugin for Fortran Feature Characterization

As new compute systems are developed, there is still a need to compile and execute codes authored in Fortran on these leading edge systems. In order to achieve this, development of compilers that support the latest hardware is continuously under development. Though the specification of Fortran is extensive, it is helpful to compiler authors to be able to prioritize the development of key features in order to get certain codes deemed important, e.g., applications of interest to leadership computing facilities, executable on leading edge compute systems. Identifying key features though is largely done through querying software experts or users of the Fortran applications of interest, who then manually report what features are and are not present. This exercise can both time consuming and error prone. To automate this process, we present a compiler plugin to Flang, the Fortran frontend for LLVM. This plugin is a tool that operates on the parse tree representation generated by Flang and detects key features based on walking parse tree nodes that correspond to features of interest. We show the result of our tool on four applications, three of which were manually profiled by software experts. We show the discrepancies between our tool and the manual characterization of the three applications, as well as generate a characterization for an application not yet profiled. We intend to open-source our tool in order to invite the community to benefit from the tool and make contributions for other features.

Cabrera, Anthony [ORNL]↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗

Softwarized Federations of Science Instruments with Edge-Continuum Containers

Significant expansion of capabilities of DOE science complex of supercomputers, instruments and networks is expected as powerful experimental facilities, exascale computers and terabit networks are added. Combined with the advances in edge and cloud computing technologies, DOE science users now have the promise of unprecedented execution of complex, continuum workflows, namely, small and latency-sensitive computations at the edge and site, massive computations at remote HPC systems, and everything in between on the cloud. But, bringing this capability to the science user requires overcoming the overwhelming complexity of forming the federations of systems, and efficiently and effectively orchestrating the workflows while ensuring high utilization of the expensive facilities. Current manual configuration of the federated systems simply will not scale, since the coordination across sites may take weeks to months, often leading to under-utilized and hard-to-diagnose compositions. A powerful, composable software stack will be developed to (i) wrap the systems so that federations can be composed fast in software, and (ii) containerize computations to be orchestrated across the edge, site, cloud and HPC resources.

Rao, Nageswara S.↗

Optimizing cloud motion estimation on the edge with phase correlation and optical flow

Abstract. Phase correlation (PC) is a well-known method for estimating cloud motion vectors (CMVs) from infrared and visible spectrum images. Commonly, phase shift is computed in the small blocks of the images using the fast Fourier transform. In this study, we investigate the performance and the stability of the blockwise PC method by changing the block size, the frame interval, and combinations of red, green, and blue (RGB) channels from the total sky imager (TSI) at the United States Atmospheric Radiation Measurement user facility's Southern Great Plains site. We find that shorter frame intervals, followed by larger block sizes, are responsible for stable estimates of the CMV, as suggested by the higher autocorrelations. The choice of RGB channels has a limited effect on the quality of CMVs, and the red and the grayscale images are marginally more reliable than the other combinations during rapidly evolving low-level clouds. The stability of CMVs was tested at different image resolutions with an implementation of the optimized algorithm on the Sage cyberinfrastructure test bed. We find that doubling the frame rate outperforms quadrupling the image resolution in achieving CMV stability. The correlations of CMVs with the wind data are significant in the range of 0.38–0.59 with a 95 % confidence interval, despite the uncertainties and limitations of both datasets. A comparison of the PC method with constructed data and the optical flow method suggests that the post-processing of the vector field has a significant effect on the quality of the CMV. The raindrop-contaminated images can be identified by the rotation of the TSI mirror in the motion field. The results of this study are critical to optimizing algorithms for edge-computing sensor systems.

54 ENVIRONMENTAL SCIENCES↗

Data Center Market Report

The data center market is poised to explode in the coming decade due to undeniable drivers such as continued adoption of generative AI, increased data storage needs, and enterprise integration of AI in numerous industries [1] [2] [3]. Scalable power and increased computational capacity are at the forefront of considerations for hyperscalers, the major cloud service providers in this space. Lawrence Livermore National Laboratory is uniquely poised to help with informed decision making for data center market leaders during this phase of explosive expansion. National grid modeling expertise and cutting edge innovations in computer cooling systems place LLNL in an enviable position for creating economic impact in the data center industry by leveraging its expertise in these areas which can help the data center market keep up with growing demand.

97 MATHEMATICS AND COMPUTING↗

A Privacy-Aware Federated Learning Framework for Distributed Energy Resource Analytics in Constrained Environments

To be resilient against extreme weather events, the rural communities in Puerto Rico are leveraging distributed energy resources (DER). However, computing frameworks sup-porting the grid in critical decision-making are still largely centralized. Sensitive consumer data are transmitted over the Internet or cellular networks to a secondary or tertiary node. It guarantees better situational awareness at the cost of a wider attack surface, jeopardizing user privacy, as more DER come online. Cloud, Edge, and Fog computing all require data aggregation at some level. This paper introduces a privacy-aware federated learning framework that leverages the Fog model by pushing analytics all the way to the DER and load assets. These local models train on individual asset data and transmit only learned parameters (such as weights) over secure communications to a global decision-maker. By abstracting personally identifiable consumer data without impacting decision optimality, this framework better aligns with distributed power generation paradigm.

Sundararajan, Aditya↗

Modeling of carbon and tungsten transient dust influx in tokamak edge plasma

The paper presents computer simulation studies of burst injection of carbon and tungsten dust particles in DIII-D-like edge plasmas. The injection causes a large transient influx of the low- and high-Z impurities associated with the dust ablation in the plasmas. The dust transport and the effects of the ablated impurities on the edge plasma dynamics in a modern mid-size tokamak geometry are investigated for low- and high-power plasma discharge conditions. The core plasma contamination with dust-ablated impurities and the factors affecting it are evaluated.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Final Reports of the 2019 Los Alamos National Laboratory Computational Physics Student Summer Workshop

The Los Alamos National Laboratory Computational Physics Workshop is intended to educate select students in problems of computational physics, while exciting them about problems of particular interest to LANL. The long term goal is to train a cadre of future researchers with strong connections to LANL and both interest and expertise in the problems LANL faces. This year’s workshop ran from June 10 – August 16, 2019 and once again attracted a phenomenal group of students, whose work is presented in the following pages. The students worked with LANL staff mentors in teams of two, doing original research. In addition, they attended a lecture series focusing on both the basics and the cutting edge challenges of computational physics (Table 1). A new item this year was a modification to our Meet LANL series of brown-bag lunch session, moving from focusing on individual researchers to panel presentations from the different groups within XCP. As always, the mentors were drawn from across disciplines and divisions. Projects looked physics ranging from quantum interactions, up to solar system development.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Learned adaptive properties for mitigation of weight perturbations in embedded spiking networks

Recent years have seen an increased importance of neural network inference in edge-based scenarios, which impose size and power constraints requiring novel computing devices. These same edge scenarios may require operating over long periods of time, or exposure to extreme environments, resulting in a drift of neural network weights that cause degraded performance. In searching for ways to develop neural network approaches that perform robustly under these conditions, we propose a biologically-inspired mechanism for the dynamic adaptation of within-neuron parameters that is guided by a global context signal carrying information about perturbations and variability in incoming stimuli. Specifically, we demonstrate that adaptive voltage thresholds or neuronal time constants, when informed by a global context signal, can enable network-level mechanisms to recover from perturbed synaptic weights. Consistent with prior literature, the context-modulated approach is effective for recurrent, but not feedforward networks, by modulating network level dynamics. We demonstrate this approach successfully recovers performance in image classification tasks and spatiotemporal tracking tasks under idealized and Gaussian noise as well as for realistic perturbations from a memristive device when exposed to ionizing radiation. Finally, we discuss how this approach enables the design of robust and energy-efficient neuromorphic systems that perform well, even in resource-constrained scenarios with extreme environments such as edge processing.

context modulation↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

High–Energy Earth–Abundant Cathodes with Enhanced Cationic/Anionic Redox for Sustainable and Long–Lasting Na–Ion Batteries

Layered iron/manganese-based oxides are a class of promising cathode materials for sustainable batteries due to their high energy densities and earth abundance. However, the stabilization of cationic and anionic redox reactions in these cathodes during cycling at high voltage remain elusive. Here, an electrochemically/thermally stable P2-Na 0.67 Fe 0.3 Mn 0.5 Mg 0.1 Ti 0.1 O 2 cathode material with zero critical elements is designed for sodium-ion batteries (NIBs) to realize a highly reversible capacity of ≈210 mAh g –1 at 20 mA g –1 and good cycling stability with a capacity retention of 74% after 300 cycles at 200 mA g –1 , even when operated with a high charge cut-off voltage of 4.5 V versus sodium metal. Combining a suite of cutting-edge characterizations and computational modeling, it is shown that Mg/Ti co-doping leads to stabilized surface/bulk structure at high voltage and high temperature, and more importantly, enhances cationic/anionic redox reaction reversibility over extended cycles with the suppression of other undesired oxygen activities. This work fundamentally deepens the failure mechanism of Fe/Mn-based layered cathodes and highlights the importance of dopant engineering to achieve high-energy and earth-abundant cathode material for sustainable and long-lasting NIBs.

25 ENERGY STORAGE↗

Selective Hydration of Rutile TiO 2 as a Strategy for Site-Selective Atomic Layer Deposition

In this report the feasibility of a site-selective hydration strategy that enables site-selective atomic layer deposition (ALD) is investigated among four rutile TiO 2 facets [(110), (100), (101) and (001)] and their most prevalent step edges. First-principles simulations of asymmetric slab models were utilized to create accurate representations of pristine terrace and step edge sites. The adsorption free energies for molecular and dissociative adsorption of H 2 O were calculated to evaluate this strategy as a viable route to step edge selectivity. We predict that selective hydroxylation is possible on the 110 and 001 step edges and further computationally evaluate three metalorganic ALD precursors for their compatibility with the selective hydration strategy. Experimental evidence for delayed nucleation of ALD on rutile (001), (110), and (100) TiO2 single crystals corroborates predictions of the dehydration of the surface and suggests the possibility of site-selective ALD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AI

As Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions.

Santos Souza, Renan↗

ASAP: Automatic Synthesis of Area-Efficient and Precision-Aware CGRAs

Coarse-grained reconfigurable accelerators (CGRAs) are a promising accelerator design choice that strikes a balance between performance and adaptability to different computing patterns across various applications domains. Designing a CGRA for a specific application domain involves enormous software/hardware engineering effort. Recent research works explore loop transformations, functional unit types, network topology, and memory size to identify optimal CGRA designs given a set of kernels from a specific application do- main. Unfortunately, the impact of functional units with different precision support has rarely been investigated. To address this gap, we propose ASAP – a hardware/software co-design framework that automatically identifies and synthesizes optimal precision-aware CGRA for a set of applications of interest. Our evaluation shows that ASAP generates specialized designs 3.2×, 4.21×, and 5.8× more efficient (in terms of performance per unit of energy or area) than non-specialized homogeneous CGRAs, for the scientific computing, embedded, and edge machine learning domains, respectively, with limited accuracy loss. Moreover, ASAP provides more efficient designs than other state-of-the-art synthesis frameworks for specialized CGRAs.

artificial intelligence↗

Tournament-Based Pretraining to Accelerate Federated Learning

Advances in hardware, proliferation of compute at the edge, and data creation at unprecedented scales have made federated learning (FL) necessary for the next leap forward in pervasive machine learning. For privacy and network reasons, large volumes of data remain stranded on endpoints located in geographically austere (or at least austere network-wise) locations. However, challenges exist to the effective use of these data. To solve the system and functional level challenges, we present an three novel variants of a serverless federated learning framework. We also present tournament-based pretraining, which we demonstrate significantly improves model performance in some experiments. Overall, these extensions to FL and our novel training method enable greater focus on science rather than ML development.

Baughman, Matt↗