Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Insights into co-pyrolysis of polyethylene terephthalate and polyamide 6 mixture through experiments, kinetic modeling and machine learning

The non-isothermal pyrolysis of polyethylene terephthalate (PET), polyamide 6 (PA6), and their mixtures was studied in a thermogravimetric analyzer at different heating rates. Temperature of maximum decomposition (T max ) decreased by 25–45 °C and 35–55 °C for the PET:PA6 mixtures (3:1, 1:1, 1:3) compared to PET and PA6, respectively. The kinetic analysis was initially carried out using isoconversional method. However, the dependency of activation energy on conversion was observed for the co-pyrolysis of PET and PA6 that suggested the occurrence of multi-step reactions in the mixtures. Distributed activation energy model (DAEM) was used in this study to describe the multistep reactions occurring during pyrolysis of PET:PA6 mixtures. Here, in this work, a four-parallel reaction DAEM was developed to describe the pyrolysis kinetics of PET:PA6 mixtures. The apparent mean activation energies (E o ) for PET, PA6, and mixtures varied in the range of 244–255, 140–215, and 138–255 kJ mol –1 , respectively. The mass loss profiles of PET and PA6 mixtures were also modeled using artificial neural network (ANN). Out of 155 ANN models, the best prediction was made by ANN511 with R 2 greater than 0.997 for both test and unseen data. The interaction effects observed through TGA experiments and subsequent kinetic analysis were further assessed in terms of product composition using analytical pyrolysis coupled with gas chromatograph/mass spectrometer (Py-GC/MS). Co-pyrolysis of PET and PA6 resulted in the formation of new aromatic compounds with nitrogen-containing functional groups, which were not detected when PET or PA6 were pyrolyzed individually.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The U.S. High-Performance Computing Consortium in the Fight Against COVID-19

U.S. computing leaders, including Department of Energy National Laboratories, have partnered with universities, government agencies, and the private sector to research responses to COVID-19, providing an unprecedented collection of resources that include some of the fastest computers in the world. For HPC users, these leadership machines will drive the AI to accelerate the discovery of promising treatments, enable at-scale simulations to understand the virus’s protein structure and attack mechanisms, and help inform policymakers to deploy resources effectively.

60 APPLIED LIFE SCIENCES↗

Classifying metal‐binding sites with neural networks

Abstract To advance our ability to predict impacts of the protein scaffold on catalysis, robust classification schemes to define features of proteins that will influence reactivity are needed. One of these features is a protein's metal‐binding ability, as metals are critical to catalytic conversion by metalloenzymes. As a step toward realizing this goal, we used convolutional neural networks (CNNs) to enable the classification of a metal cofactor binding pocket within a protein scaffold. CNNs enable images to be classified based on multiple levels of detail in the image, from edges and corners to entire objects, and can provide rapid classification. First, six CNN models were fine‐tuned to classify the 20 standard amino acids to choose a performant model for amino acid classification. This model was then trained in two parallel efforts: to classify a 2D image of the environment within a given radius of the central metal binding site, either an Fe ion or a [2Fe‐2S] cofactor, with the metal visible (effort 1) or the metal hidden (effort 2). We further used two sub‐classifications of the [2Fe‐2S] cofactor: (1) a standard [2Fe‐2S] cofactor and (2) a Rieske [2Fe‐2S] cofactor. The accuracy for the model correctly identifying all three defined features was >95%, despite our perception of the increased challenge of the metalloenzyme identification. This demonstrates that machine learning methodology to classify and distinguish similar metal‐binding sites, even in the absence of a visible cofactor, is indeed possible and offers an additional tool for metal‐binding site identification in proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Massively scalable Kerr comb-driven silicon photonic link

Abstract The growth of computing needs for artificial intelligence and machine learning is critically challenging data communications in today’s data-centre systems. Data movement, dominated by energy costs and limited ‘chip-escape’ bandwidth densities, is perhaps the singular factor determining the scalability of future systems. Using light to send information between compute nodes in such systems can dramatically increase the available bandwidth while simultaneously decreasing energy consumption. Through wavelength-division multiplexing with chip-based microresonator Kerr frequency combs, independent information channels can be encoded onto many distinct colours of light in the same optical fibre for massively parallel data transmission with low energy. Although previous high-bandwidth demonstrations have relied on benchtop equipment for filtering and modulating Kerr comb wavelength channels, data-centre interconnects require a compact on-chip form factor for these operations. Here we demonstrate a massively scalable chip-based silicon photonic data link using a Kerr comb source enabled by a new link architecture and experimentally show aggregate single-fibre data transmission of 512 Gb s −1 across 32 independent wavelength channels. The demonstrated architecture is fundamentally scalable to hundreds of wavelength channels, enabling massively parallel terabit-scale optical interconnects for future green hyperscale data centres.

Rizzo, Anthony (ORCID:000000034752797X)↗

AI-optimized detector design for the future Electron-Ion Collider: the dual-radiator RICH case

Advanced detector R&D requires performing computationally intensive and detailed simulations as part of the detector-design optimization process. Here, we propose a general approach to this process based on Bayesian optimization and machine learning that encodes detector requirements. As a case study, we focus on the design of the dual-radiator Ring Imaging Cherenkov (dRICH) detector under development as a potential component of the particle-identification system at the future Electron-Ion Collider (EIC). The EIC is a US-led frontier accelerator project for nuclear physics, which has been proposed to further explore the structure and interactions of nuclear matter at the scale of sea quarks and gluons. We show that the detector design obtained with our automated and highly parallelized framework outperforms the baseline dRICH design within the assumptions of the current model. Our technique can be applied to any detector R&D, provided that realistic simulations are available.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Future of plasma etching for microelectronics: Challenges and opportunities

Plasma etching is an essential semiconductor manufacturing technology required to enable the current microelectronics industry. Along with lithographic patterning, thin-film formation methods, and others, plasma etching has dynamically evolved to meet the exponentially growing demands of the microelectronics industry that enables modern society. At this time, plasma etching faces a period of unprecedented changes owing to numerous factors, including aggressive transition to three-dimensional (3D) device architectures, process precision approaching atomic-scale critical dimensions, introduction of new materials, fundamental silicon device limits, and parallel evolution of post-CMOS approaches. The vast growth of the microelectronics industry has emphasized its role in addressing major societal challenges, including questions on the sustainability of the associated energy use, semiconductor manufacturing related emissions of greenhouse gases, and others. The goal of this article is to help both define the challenges for plasma etching and point out effective plasma etching technology options that may play essential roles in defining microelectronics manufacturing in the future. The challenges are accompanied by significant new opportunities, including integrating experiments with various computational approaches such as machine learning/artificial intelligence and progress in computational approaches, including the realization of digital twins of physical etch chambers through hybrid/coupled models. These prospects can enable innovative solutions to problems that were not available during the past 50 years of plasma etch development in the microelectronics industry. To elaborate on these perspectives, the present article brings together the views of various experts on the different topics that will shape plasma etching for microelectronics manufacturing of the future.

Engineering↗

High energy density picoliter-scale zinc-air microbatteries for colloidal robotics

The recent interest in microscopic autonomous systems, including microrobots, colloidal state machines, and smart dust, has created a need for microscale energy storage and harvesting. However, macroscopic materials for energy storage have noted incompatibilities with microfabrication techniques, creating substantial challenges to realizing microscale energy systems. Here, we photolithographically patterned a microscale zinc/platinum/SU-8 system to generate the highest energy density microbattery at the picoliter (10 −12 liter) scale. The device scavenges ambient or solution-dissolved oxygen for a zinc oxidation reaction, achieving an energy density ranging from 760 to 1070 watt-hours per liter at scales below 100 micrometers lateral and 2 micrometers thickness in size. The parallel nature of photolithography processes allows 10,000 devices per wafer to be released into solution as colloids with energy stored on board. Within a volume of only 2 picoliters each, these primary microbatteries can deliver open circuit voltages of 1.05 ± 0.12 volts, with total energies ranging from 5.5 ± 0.3 to 7.7 ± 1.0 microjoules and a maximum power near 2.7 nanowatts. We demonstrated that such systems can reliably power a micrometer-sized memristor circuit, providing access to nonvolatile memory. We also cycled power to drive the reversible bending of microscale bimorph actuators at 0.05 hertz for mechanical functions of colloidal robots. Additional capabilities, such as powering two distinct nanosensor types and a clock circuit, were also demonstrated. The high energy density, low volume, and simple configuration promise the mass fabrication and adoption of such picoliter zinc-air batteries for micrometer-scale, colloidal robotics with autonomous functions.

Robotics↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

Design and Construction of a High-Resolution Hodoscope for the GlueX Experiment with High-Statistics Analysis of the p0, ¿, and ¿1 Photoproduction Cross Sections from the RadPhi Experiment

Differential cross sections for forward-angle photoproduction of p0, ¿, and ¿ 1 pseudoscalar mesons were measured using data from the RadPhi experiment conducted in Hall B at Jef ferson Lab. RadPhi utilized a tagged bremsstrahlung photon beam incident on a stationary 9Be target, with a detector system configured to trigger on a recoil proton in coincidence with multiple neutral showers in the calorimeter. Events were reconstructed and subjected to kinematic constraints, with background suppressed via sideband subtraction guided by Monte Carlo modeling of background contributions. Cross sections were extracted over the photon energy range 4.4– 5.4 GeV and binned in invariant momentum transfer t, providing measurements from one of the first high-statistics experiments of forward ¿ and ¿1 pro duction from a nuclear target at these energies. Acceptance corrections were applied using a detailed GEANT-based simulation of the detector geometry and response. The resulting cross sections are consistent with 2020 CLAS results, when scaled by the number of protons in beryllium, and show broad agreement with other data and theoretical models. In parallel, a high-resolution photon tagger detector, the Tagger Microscope (TAGM), was designed, constructed, and commissioned for the GlueX experiment in Hall D at Jefferson Lab. The TAGM was developed to provide high-rate tagging capability in the coherent bremsstrahlung peak by detecting post-bremsstrahlung electrons across a one GeV range along the focal plane of the tagging spectrometer. The detector consists of a 5ˆ102 array of 2ˆ2 mm2 square BCF-20 plastic scintillating fibers thermally fused to BCF-98 light guide fibers optically coupled to silicon photomultipliers. These fibers are mounted in a precision machined framework enabling fine positional adjustments to maintain precise alignment with post-bremsstrahlung electron trajectories, while ensuring mechanical rigidity, thermal stability, optical isolation, minimal inactive area, and radiation shielding for electronics. The construction effort involved extensive testing of fiber quality, light transmission, thermal fusing, radiation hardness, and defect analysis using SEM and EDX techniques. Following its installation and commissioning, the TAGM became a critical component of the GlueX beamline, enabling high-rate tagging essential for studies of hybrid mesons and gluonic ex citations.

McIntyre, James [Univ. of Connecticut, Storrs, CT ↗

Data-flow parallelism for high-energy and nuclear physics frameworks

The processing tasks of an event-processing workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this talk, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. Building on the Meld project as presented at CHEP2023, we demonstrate that all common processing idioms supported by current frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Synchronous High-frequency Distributed Readout For Edge Processing At The Fermilab Main Injector And Recycler

The Main Injector (MI) was commissioned using data acquisition systems developed for the Fermilab Main Ring in the 1980s. New VME-based instrumentation was commissioned in 2006 for beam loss monitors (BLM)[2], which provided a more systematic study of the machine and improved displays of routine operation. However, current projects are demanding more data and at a faster rate from this aging hardware. One such project, Real-time Edge AI for Distributed Systems (READS), requires the high-frequency, low-latency collection of synchronized BLM readings from around the approximately two-mile accelerator complex. Significant work has been done to develop new hardware to monitor the VME backplane and broadcast BLM measurements over Ethernet, while not disrupting the existing operations critical functions of the BLM system. This paper will detail the design, implementation, and testing of this parallel data pathway.

43 PARTICLE ACCELERATORS↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗

Genetic algorithm optimization of nuclear criticality experiment for reduction of intermediate-energy 239 Pu nuclear data uncertainties

Nuclear criticality experiments are conducted to investigate specific nuclear data important for safe handling and storage of fissile materials, reactor design and operation, and the validation of radiation transport codes. Incorrect or uncertain nuclear data can prohibitively impact operational safety limits, reactor licensing, and predictive simulation capability; therefore, integral measurements from criticality experiments are necessary and should be performed frequently. To maximize the impact of the integral measurements, it is important to consider experiment geometry, material selection, and component dimensions. When taking these considerations into account, the experiment design process becomes iterative and very time intensive. This work utilizes a genetic algorithm to efficiently explore potential nuclear criticality experiment designs for the Laboratory Directed Research & Development project PARADIGM (PARallel Approach of Differential and InteGral Measurements) at Los Alamos National Laboratory. In this paper, the building blocks of the genetic algorithm are discussed in detail, the genetic algorithm methodology is verified, and the genetic algorithm is used to produce three candidate experiment models for the final PARADIGM design. The three candidate models produced by the genetic algorithm consist of copper-reflected assemblies containing 14 repeating units of alumina, graphite, boron, and plutonium plates. Furthermore, in addition to the optimization results, final design considerations are also discussed for designs with a height and/or weight very close to or slightly above assembly machine operational limits.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Toward fully automated UED operation using two-stage machine learning model

To demonstrate the feasibility of automating UED operation and diagnosing the machine performance in real time, a two-stage machine learning (ML) model based on self-consistent start-to-end simulations has been implemented. This model will not only provide the machine parameters with adequate precision, toward the full automation of the UED instrument, but also make real-time electron beam information available as single-shot nondestructive diagnostics. Furthermore, based on a deep understanding of the root connection between the electron beam properties and the features of Bragg-diffraction patterns, we have applied the hidden symmetry as model constraints, successfully improving the accuracy of energy spread prediction by a factor of five and making the beam divergence prediction two times faster. The capability enabled by the global optimization via ML provides us with better opportunities for discoveries using near-parallel, bright, and ultrafast electron beams for single-shot imaging. It also enables directly visualizing the dynamics of defects and nanostructured materials, which is impossible using present electron-beam technologies.

36 MATERIALS SCIENCE↗

Measurement-Based Quantum Thermal Machines with Feedback Control

We investigated coupled-qubit-based thermal machines powered by quantum measurements and feedback. We considered two different versions of the machine: (1) a quantum Maxwell’s demon, where the coupled-qubit system is connected to a detachable single shared bath, and (2) a measurement-assisted refrigerator, where the coupled-qubit system is in contact with a hot and cold bath. In the quantum Maxwell’s demon case, we discuss both discrete and continuous measurements. We found that the power output from a single qubit-based device can be improved by coupling it to the second qubit. We further found that the simultaneous measurement of both qubits can produce higher net heat extraction compared to two setups operated in parallel where only single-qubit measurements are performed. In the refrigerator case, we used continuous measurement and unitary operations to power the coupled-qubit-based refrigerator. We found that the cooling power of a refrigerator operated with swap operations can be enhanced by performing suitable measurements.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Design and Validation of a High-Throughput Reductive Catalytic Fractionation Method

Reductive catalytic fractionation (RCF) is a promising method to extract and depolymerize lignin from biomass, and bench-scale studies have enabled considerable progress in the past decade. RCF experiments are typically conducted in pressurized batch reactors with volumes ranging between 50 and 1000 mL, limiting the throughput of these experiments to one to six reactions per day for an individual researcher. Here, we report a high-throughput RCF (HTP-RCF) method in which batch RCF reactions are conducted in 1 mL wells machined directly into Hastelloy reactor plates. The plate reactors can seal high pressures produced by organic solvents by vertically stacking multiple reactor plates, leading to a compact and modular system capable of performing 240 reactions per experiment. Using this setup, we screened solvent mixtures and catalyst loadings for hydrogen-free RCF using 50 mg poplar and 0.5 mL reaction solvent. The system of 1:1 isopropanol/methanol showed optimal monomer yields and selectivity to 4-propyl substituted monomers, and validation reactions using 75 mL batch reactors produced identical monomer yields. To accommodate the low material loadings, we then developed a workup procedure for parallel filtration, washing, and drying of samples and a 1H nuclear magnetic resonance spectroscopy method to measure the RCF oil yield without performing liquid-liquid extraction. As a demonstration of this experimental pipeline, 50 unique switchgrass samples were screened in RCF reactions in the HTP-RCF system, revealing a wide range of monomer yields (21-36%), S/G ratios (0.41-0.93), and oil yields (40-75%). These results were successfully validated by repeating RCF reactions in 75 mL batch reactors for a subset of samples. We anticipate that this approach can be used to rapidly screen substrates, catalysts, and reaction conditions in high-pressure batch reactions with higher throughput than standard batch reactors.

BIOMASS FUELS,INORGANIC, ORGANIC, PHYSICAL, AND AN↗

Deep Learning Approaches to Surrogates for Solving the Diffusion Equation for Mechanistic Real-World Simulations

In many mechanistic medical, biological, physical, and engineered spatiotemporal dynamic models the numerical solution of partial differential equations (PDEs), especially for diffusion, fluid flow and mechanical relaxation, can make simulations impractically slow. Biological models of tissues and organs often require the simultaneous calculation of the spatial variation of concentration of dozens of diffusing chemical species. One clinical example where rapid calculation of a diffusing field is of use is the estimation of oxygen gradients in the retina, based on imaging of the retinal vasculature, to guide surgical interventions in diabetic retinopathy. Furthermore, the ability to predict blood perfusion and oxygenation may one day guide clinical interventions in diverse settings, i.e., from stent placement in treating heart disease to BOLD fMRI interpretation in evaluating cognitive function (Xie et al., 2019; Lee et al., 2020). Since the quasi-steady-state solutions required for fast-diffusing chemical species like oxygen are particularly computationally costly, we consider the use of a neural network to provide an approximate solution to the steady-state diffusion equation. Machine learning surrogates, neural networks trained to provide approximate solutions to such complicated numerical problems, can often provide speed-ups of several orders of magnitude compared to direct calculation. Surrogates of PDEs could enable use of larger and more detailed models than are possible with direct calculation and can make including such simulations in real-time or near-real time workflows practical. Creating a surrogate requires running the direct calculation tens of thousands of times to generate training data and then training the neural network, both of which are computationally expensive. Often the practical applications of such models require thousands to millions of replica simulations, for example for parameter identification and uncertainty quantification, each of which gains speed from surrogate use and rapidly recovers the up-front costs of surrogate generation. We use a Convolutional Neural Network to approximate the stationary solution to the diffusion equation in the case of two equal-diameter, circular, constant-value sources located at random positions in a two-dimensional square domain with absorbing boundary conditions. Such a configuration caricatures the chemical concentration field of a fast-diffusing species like oxygen in a tissue with two parallel blood vessels in a cross section perpendicular to the two blood vessels. To improve convergence during training, we apply a training approach that uses roll-back to reject stochastic changes to the network that increase the loss function. The trained neural network approximation is about 1000 times faster than the direct calculation for individual replicas. Because different applications will have different criteria for acceptable approximation accuracy, we discuss a variety of loss functions and accuracy estimators that can help select the best network for a particular application. We briefly discuss some of the issues we encountered with overfitting, mismapping of the field values and the geometrical conditions that lead to large absolute and relative errors in the approximate solution.

60 APPLIED LIFE SCIENCES↗

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)↗