Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “resource consumption”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Quantum Computing in the Cloud: Analyzing job and machine characteristics

As the popularity of quantum computing continues to grow, quantum machine access over the cloud is critical to both academic and industry researchers across the globe. And as cloud quantum computing demands increase exponentially, the analysis of resource consumption and execution characteristics are key to efficient management of jobs and resources at both the vendor-end as well as the client-end. While the analysis of resource consumption and management are popular in the classical HPC domain, it is severely lacking for more nascent technology like quantum computing. This paper is a first-of-its-kind academic study, analyzing various trends in job execution and resources consumption / utilization on quantum cloud systems. We focus on IBM Furthermore, quantum systems and analyze characteristics over a two year period, encompassing over 6000 jobs which contain over 600,000 quantum circuit executions and correspond to almost 10 billion “shots” or trials over 20+ quantum machines. Specifically, we analyze trends focused on, but not limited to, execution times on quantum machines, queuing/waiting times in the cloud, circuit compilation times, machine utilization, as well as the impact of job and machine characteristics on all of these trends. Furthermore, our analysis identifies several similarities and differences with classical HPC cloud systems. Based on our insights, we make recommendations and contributions to improve the management of resources and jobs on future quantum cloud systems.

42 ENGINEERING↗

wa-hls4ml: A GNN Surrogate Model for hls4ml

Recent advancements in use of machine learning techniques on field-programmable gate arrays (FPGAs) have allowed for implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area must be strictly bounded. The hls4ml framework is a procedure for converting from trained machine learning model software, to a synthesis result that can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it is possible that the model is unable to be converted into a synthesis result, or that the resource consumption of the model will exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model which uses a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of an arbitrary model when passed through the hls4ml procedure, without the time consumption of actually running the pipeline.

43 PARTICLE ACCELERATORS↗

A Graph Neural Network Surrogate Model for hls4ml

Recent advancements in use of machine learning (ML) techniques on field-programmable gate arrays (FPGAs) have allowed for the implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area are strictly bounded. The hls4ml framework is a procedure that converts trained ML model software to a synthesis result to can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it may not be possible to successfully convert a model into a synthesis result, or the resource consumption of the model may exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model using a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of a given model when passed through the hls4ml pipeline, without needing to run the pipeline.

Plotnikov, Dennis↗

Adaptive job and resource management for the growing quantum cloud

As the popularity of quantum computing continues to grow, efficient quantum machine access over the cloud is critical to both academic and industry researchers across the globe. And as cloud quantum computing demands increase exponentially, the analysis of resource consumption and execution characteristics are key to efficient management of jobs and resources at both the vendor-end as well as the client-end. While the analysis and optimization of job / resource consumption and management are popular in the classical HPC domain, it is severely lacking for more nascent technology like quantum computing.This paper proposes optimized adaptive job scheduling to the quantum cloud taking note of primary characteristics such as queuing times and fidelity trends across machines, as well as other characteristics such as quality of service guarantees and machine calibration constraints. Key components of the proposal include a) a prediction model which predicts fidelity trends across machine based on compiled circuit features such as circuit depth and different forms of errors, as well as b) queuing time prediction for each machine based on execution time estimations. Altogether, this proposal is evaluated on simulated IBM machines across a diverse set of quantum applications and system loading scenarios, and is able to reduce wait times by over 3x and improve fidelity by over 40% on specific usecases, when compared to traditional job schedulers.

97 MATHEMATICS AND COMPUTING↗

Characteristics of U.S. Energy Production using Nuclear Fission

When considering potential energy production technologies for the future, a critical consideration centers on the question of “how green” the technology is. Here, the word “green” implies that the technology has a zero or minimal impact on the public and environment while still providing benefits by way of electricity, heat, and other products such as hydrogen and water. And, these potential impacts must be considered over the lifecycle of the technology deployment, from design, construction, operation, and disposition. Ideally, a green energy technology would be net-zero (i.e., having no to almost negligible contribution) on five key elements: 1. Greenhouse gas emissions, including carbon dioxide (CO 2 ), methane, nitrous oxide, and fluorinated gases such as ozone-depleting gases 2. Water consumption 3. Material resource consumption 4. Disposition of wastes 5. Public, flora, and fauna safety including deaths, health, and environmental impacts. Society would benefit from net-zero impacts of the five elements above while having low-cost energy. As society starts to replace fossil fuels in the energy mix, we need to consider the possible impacts of adopted technologies. We need a production approach that provides large quantities of energy while being safe, reliable, economical, and sustainable. In this report, we describe a variety of characteristics related to the five elements above to provide a fact-supported, science-based depiction of nuclear fission as an energy providing technology. By better understanding how fission power is nearly net-zero in the five elements above, we can position our thinking to align with the overarching goal to provide society with a low-impact energy source.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Efficient and choreographed quality-of- service management in dense 6G verticals with high-speed mobility requirements

Future 6G networks are envisioned to support very heterogeneous and extreme applications (known as verticals). Some examples are further-enhanced mobile broadband communications, where bitrates could go above one terabit per second, or extremely reliable and low-latency communications, whose end-to-end delay must be below one hundred microseconds. To achieve that ultra-high Quality-of-Service, 6G networks are commonly provided with redundant resources and intelligent management mechanisms to ensure that all devices get the expected performance. But this approach is not feasible or scalable for all verticals. Specifically, in 6G scenarios, mobile devices are expected to have speeds greater than 500 kilometers per hour, and device density will exceed ten million devices per square kilometer. In those verticals, resources cannot be redundant as, because of such a huge number of devices, Quality-of-Service requirements are pushing the effective performance of technologies at physical level. And, on the other hand, high-speed mobility prevents intelligent mechanisms to be useful, as devices move around and evolve faster than the usual convergence time of those intelligent solutions. New technologies are needed to fill this unexplored gap. Therefore, in this paper we propose a choreographed Quality-of-Service management solution, where 6G base stations predict the evolution of verticals at real-time, and run a lightweight distributed optimization algorithm in advance, so they can manage the resource consumption and ensure all devices get the required Quality-of-Service. Prediction mechanism includes mobility models (Markov, Bayesian, etc.) and models for time-variant communication channels. Besides, a traffic prediction solution is also considered to explore the achieved Quality-of-Service in advance. The optimization algorithm calculates an efficient resource distribution according to the predicted future vertical situation, so devices achieve the expected Quality-of-Service according to the proposed traffic models. An experimental validation based on simulation tools is also provided. Results show that the proposed approach reduces up to 12% of the network resource consumption for a given Quality-of-Service.

42 ENGINEERING↗

Critical Literature Review of Quantitative Sustainability Assessment Methods for the Circular Economy

The circular economy (CE) has been proposed to be an operational framework for sustainable development that decouples economic growth from resource consumption. The CE ambition is to maximize the retention of value in products, materials, and resources in the economy over time with the help of CE strategies such as selling a service rather than a product, reusing and repairing products or their components, and recycling. The social, economic, and environmental performances of CE strategies need to be measured against their linear counterparts to avoid strategies that increase circularity but have other unintended externalities. However, there is currently no tool specifically designed to compare circular to linear systems and, thus, various methods from different fields have been applied. This session aims at reviewing, contrasting, and critiquing different methods that have been applied to assess the CE until now, along with an up to date state of the science in this field. Methods from the industrial ecology field have most often been applied to study the CE. However, the transition to CE involves both technological improvements as well as social changes. New business models such as collaborative consumption models (e.g., Uber or Airbnb) are examples of the new patterns of production and consumption of the CE, which may be difficult to analyze from a purely industrial ecology perspective. Methods from complexity science and humanities could therefore complement the industrial ecology perspective to extend the scope of the analysis. Such an approach could also answer a longstanding criticism of the CE which, in contrast with sustainability, solely focuses on the environment and the economy. Moreover, the hybridization of two or more existing methods can yield additional capabilities, which may enable the exploration of additional CE-related research questions. The 90 minutes session will have several presentations and conclude with a moderated, interactive panel discussion on how methods from different fields could be combined to harness their relative strengths.

circular economy↗

HGQ: High Granularity Quantization for Real-time Neural Networks on FPGAs

Neural networks with sub-microsecond inference latency are required by many critical applications. Targeting such applications deployed on FPGAs, we present High Granularity Quantization (HGQ), a quantization-aware training framework that optimizes parameter bit-widths through gradient descent. Unlike conventional methods, HGQ determines the optimal bit-width for each parameter independently, making it suitable for hardware platforms supporting heterogeneous arbitrary precision arithmetic. In our experiments, HGQ shows superior performance compared to existing network compression methods, achieving orders of magnitude reduction in resource consumption and latency while maintaining the accuracy on several benchmark tasks. These improvements enable the deployment of complex models previously infeasible due to resource or latency constraints. HGQ is open-source and is used for developing next-generation trigger systems at the CERN ATLAS and CMS experiments for particle physics, enabling the use of advanced machine learning models for real-time data selection with sub-microsecond latency.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SymbolNet: neural symbolic regression with adaptive dynamic pruning for compression

Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).

Tsoi, Ho Fung (ORCID:0000000225502184)↗

Ultrafast jet classification at the HL-LHC

Abstract Three machine learning models are used to perform jet origin classification. These models are optimized for deployment on a field-programmable gate array device. In this context, we demonstrate how latency and resource consumption scale with the input size and choice of algorithm. Moreover, the models proposed here are designed to work on the type of data and under the foreseen conditions at the CERN large hadron collider during its high-luminosity phase. Through quantization-aware training and efficient synthetization for a specific field programmable gate array, we show that O ( 100 ) ns inference of complex architectures such as Deep Sets and Interaction Networks is feasible at a relatively low computational resource cost.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

An innovative approach for atrazine electrochemical oxidation modelling: Process parameter effect, intermediate formation and kinetic constant assessment

Water reuse for irrigation activities is becoming a crucial worldwide challenge due to the depletion of water sources. Anyway, agricultural drainage can potentially contain dangerous contaminants such as metals, pesticides, and herbicides, including atrazine. To address the need for agriculture wastewater purification, we investigated atrazine removal from simulated wastewater by electro-oxidation using platinum-coated titanium electrodes on a lab-scale experimental apparatus. The effects of electrolyte composition and concentration, i.e. ionic strength and applied current density on atrazine removal, were investigated. The results demonstrated that the electrochemical oxidation of the herbicide occurred through two routes, depending on the presence or absence of oxidizing chlorine species. The generation of intermediates during the treatment was monitored and quantified by evaluating the effect of an inert electrolyte (NaClO 4 ) versus an oxidizable chlorine species (NaCl). In both experimental conditions, five intermediates were identified, including desethyl-atrazine (DEA), hydroxyatrazine (ATZ-OH), desisopropyl-atrazine (DIA) and desethyl-desisopropyl-atrazine (DEDIA). A degradation mechanism and a model for describing hydroxyl radicals and active chlorine species contributions at ATZ oxidation were also proposed. Intermediate evolution profiles suggest that ATZ degradation can be considered as a series–parallel reaction system. Finally, the energy requirement assessment for ATZ removal was carried out. The highest ATZ removal (≅98%) was achieved with NaCl = 0.08 M, J = 60 A/m −2 , and E C = 5.83 kWh m −3 . Results highlight that atrazine removal was improved when an active chlorine species (NaCl) was present in the water solution. Moreover, the addition of chlorine species during electro-oxidation is an energy-saving strategy. Collectively, electro-oxidation technique can be efficiently applied to treat polluted water in order to meet the needs of recycling water quality and reduce resource consumption.

Electro-chemical oxidation↗

JUST-R metrics for considering energy justice in early-stage energy research

We report achieving sustainable decarbonization of the energy sector requires implementing and improving energy technologies while simultaneously managing sources of social inequity in the energy system. Centering energy justice, which has "the goal of achieving equity in both the social and economic participation in the energy system, while also remediating social, economic, and health burdens on those historically harmed by the energy system," in the transition to clean energy has become an increasingly urgent priority for social scientists, policymakers, and community activists alike. However, late-stage consideration of social impacts of energy technologies may result in identifying inequities only after substantial time, money, and effort have been expended on research and development (R&D). This issue is exemplified by concerns over environmental and human health impacts related to cobalt in lithium-ion batteries, which has spurred research into alternatives only after decades of R&D and the establishment of supply chains, infrastructure, and markets for cobalt-containing chemistries. Other examples include issues with land use and resource consumption related to first-generation biofuel feedstocks as well as occupational hazards and pollution associated with photovoltaics manufacturing. In all these cases, subsequent R&D to improve technologies or processes cannot undo the effects already experienced. Incorporating energy justice from the earliest stage of R&D will enable more just technology implementation, but integrating justice considerations into early-stage research is a challenge due to a lack of tools to assess and manage them. To fill this gap, we center early-stage research to develop the Justice Underpinning Science and Technology Research (JUST-R) metrics framework - energy justice metrics specifically targeted at early-stage researchers to assess their work on an immediate timescale. By applying these metrics to a case study focused on materials for next-generation photovoltaics, we highlight potential benefits and barriers to implementing this framework in early-stage research and discuss necessary institutional and individual actions needed for researchers to effectively leverage the tool to incorporate justice-focused criteria into R&D decision making.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber↗

Historical and Future Learning for the New Era of Multi-Terawatt Photovoltaics

Solar photovoltaics (PV) is entering a new era of multi-terawatt deployment, with 2 TW already in service and more than 75 TW predicted in many scenarios by 2050. This next era has been enabled by over five decades of cumulative advances in PV module cost reduction, performance and reliability. The current scale of deployment also introduces new needs, opportunities and challenges. In this Perspective we frame a path forwards based on learning, broadly defined as a combination of expansion of knowledge and advances through research and development, experience and collaboration. We discuss historical topics where learning has driven PV deployment until now, and emerging areas that are required to sustain high levels of future deployment. We expect progress to continue in terms of module price, performance and reliability, driven by advances in PV cell and module design, the emergence of tandem devices and increased focus on extending module lifetimes. Large-scale deployment also means large-scale sustainability and responsibility. We therefore posit that additional metrics, such as the impact on global CO2 emissions, resource consumption and design for reuse and recycling, will become increasingly important to the PV industry and provide opportunities for further learning.

14 SOLAR ENERGY↗

Fast convolutional neural networks on FPGAs with hls4ml

We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Workflow Provenance in the Computing Continuum for Responsible, Trustworthy, and Energy-Efficient AI

As Artificial Intelligence (AI) becomes more pervasive in our society, it is crucial to develop, deploy, and assess Responsible and Trustworthy AI (RTAI) models, i.e., those that consider not only accuracy but also other aspects, such as explainability, fairness, and energy efficiency. Workflow provenance data have historically enabled critical capabilities towards RTAI. Provenance data derivation paths contribute to responsible workflows through transparency in tracking artifacts and resource consumption. Provenance data are well-known for their trustworthiness helping explainability, reproducibility, and accountability. However, there are complex challenges to achieve RTAI, which are further complicated by the heterogeneous infrastructure in the computing continuum (Edge-Cloud-HPC) used to develop and deploy models. As a result, a significant research and development gap remains between workflow provenance data management and RTAI. In this paper, we present a vision of the pivotal role of workflow provenance in supporting RTAI and discuss related challenges. We present a schematic view between RTAI and provenance, and highlight open research directions.

Santos Souza, Renan↗

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗