Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multiple iterations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Neutron diffusion calculation in heterogeneous geometry based on local/global iteration using proper orthogonal decomposition

This study newly proposes a heterogeneous core calculation method based on local/global iteration using proper orthogonal decomposition (POD). By using the singular value decomposition (SVD) and the low-rank approximation, appropriate POD bases for expanding the neutron flux can be obtained from snapshot data of the neutron flux obtained by fine mesh calculations. By projection using the POD bases, the dimension of the target equation (e.g., discretized neutron diffusion equation) can be dramatically reduced. In the proposed method, POD is effectively applied to each single assembly calculation (local calculation). Furthermore, using the local/global iteration, the effective neutron multiplication factor and the neutron flux distribution in the whole core geometry can be obtained by combining the numerical results of the local calculation for each fuel assembly and the global calculation for the whole core. As a feasibility study, the proposed method is applied to a one-dimensional heterogeneous core analysis, and the accuracy is investigated by changing the total number of POD bases. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Real space iterative reconstruction for vector tomography (RESIRE-V)

Tomography has had an important impact on the physical, biological, and medical sciences. To date, most tomographic applications have been focused on 3D scalar reconstructions. However, in some crucial applications, vector tomography is required to reconstruct 3D vector fields such as the electric and magnetic fields. Over the years, several vector tomography methods have been developed. Here, we present the mathematical foundation and algorithmic implementation of REal Space Iterative REconstruction for Vector tomography, termed RESIRE-V. RESIRE-V uses multiple tilt series of projections and iterates between the projections and a 3D reconstruction. Each iteration consists of a forward step using the Radon transform and a backward step using its transpose, then updates the object via gradient descent. Incorporating with a 3D support constraint, the algorithm iteratively minimizes an error metric, defined as the difference between the measured and calculated projections. The algorithm can also be used to refine the tilt angles and further improve the 3D reconstruction. To validate RESIRE-V, we first apply it to a simulated data set of the 3D magnetization vector field, consisting of two orthogonal tilt series, each with a missing wedge. Our quantitative analysis shows that the three components of the reconstructed magnetization vector field agree well with the ground-truth counterparts. We then use RESIRE-V to reconstruct the 3D magnetization vector field of a ferromagnetic meta-lattice consisting of three tilt series. Our 3D vector reconstruction reveals the existence of topological magnetic defects with positive and negative charges. We expect that RESIRE-V can be incorporated into different imaging modalities as a general vector tomography method. To make the algorithm accessible to a broad user community, we have made our RESIRE-V MATLAB source codes and the data freely available at https://github.com/minhpham0309/RESIRE-V.

47 OTHER INSTRUMENTATION↗

Distributed Quantum Learning with co-Management in a Multi-tenant Quantum System

The rapid advancement of quantum computing has pushed classical designs into the quantum domain, breaking physical boundaries for computing-intensive and data-hungry applications with the hope that some systems may provide a quantum speedup. For example, variational quantum algorithms have been proposed for quantum neural networks to train deep learning models on qubits, achieving promising results. Existing quantum learning architectures and systems rely on single, monolithic quantum machines with abundant and stable resources, such as qubits. However, fabricating a large, monolithic quantum device is considerably more challenging than producing an array of smaller devices. In this paper, we investigate a distributed quantum system that combines multiple quantum machines into a unified system. We propose DQuLearn, which divides a quantum learning task into multiple subtasks. Each subtask can be executed distributively on individual quantum machines, with the results looping back to classical machines for subsequent training iterations. Additionally, our system supports multiple concurrent clients and dynamically manages their circuits according to the runtime status of quantum workers. Through extensive experiments, we demonstrate that DQuLearn achieves similar accuracies with significant runtime reduction, by up to 68.7% and an increase per-second circuit processing speed, by up to 3.99 times, in a 4-worker multi-tenant setting.

quantum computing↗

Sparse matrix‐vector and matrix‐multivector products for the truncated SVD on graphics processors

Summary Many practical algorithms for numerical rank computations implement an iterative procedure that involves repeated multiplications of a vector, or a collection of vectors, with both a sparse matrix and its transpose. Unfortunately, the realization of these sparse products on current high performance libraries often deliver much lower arithmetic throughput when the matrix involved in the product is transposed. In this work, we propose a hybrid sparse matrix layout, named CSRC, that combines the flexibility of some well‐known sparse formats to offer a number of appealing properties: (1) CSRC can be obtained at low cost from the popular CSR (compressed sparse row) format; (2) CSRC has similar storage requirements as CSR; and especially, (3) the implementation of the sparse product kernels delivers high performance for both the direct product and its transposed variant on modern graphics accelerators thanks to a significant reduction of atomic operations compared to a conventional implementation based on CSR. This solution thus renders considerably higher performance when integrated into an iterative algorithm for the truncated singular value decomposition (SVD), such as the randomized SVD or, as demonstrated in the experimental results, the block Golub–Kahan–Lanczos algorithm.

Aliaga, José I.↗

Challenges of conventional iterative all-atom and coarse-grained multiscale molecular dynamics

In this work, we evaluate the biomolecular dynamics behaviors when conventionally iterating between all-atom (AA) and coarse-grained (CG) molecular dynamics (MD) simulations over multiple cycles. We implemented the workflow to iterate between AA and CG in OpenMM, namely the iterative multiscale MD (iMMD) simulation workflow. In particular, we aim to identify practical applications for iterating between AA and CG simulations in a conventional manner without any constraints or model modifications. We evaluate the iMMD workflow on four representative systems, spanning folding of two soluble proteins and protein-protein as well as protein-lipid interactions of two membrane proteins. We observe that iteration between AA and CG representations could help the soluble proteins exit undesirable metastable states to fold, resulting from random protein structural distortions due to cycling. Consequently, the most reliable use of iterative AA and CG simulations appears to be to accelerating complex lipid mixing for membrane-bound protein systems rather than sampling protein conformational space. Our work explores the practical usages and limitations for iterative AA and CG simulations using readily available AA and CG force fields. The evaluated iMMD workflow in OpenMM is made available at https://github.com/lanl/iMMD.

59 BASIC BIOLOGICAL SCIENCES↗

Mesh-based multiphysics coupling acceleration for fusion neutronics through clustering for fusion blanket applications

Accurate modeling of particle transport within fusion blankets is essential for predicting performance metrics such as heat deposition and the tritium breeding ratio (TBR). However, high-fidelity coupling of thermal fluids from computational fluid dynamics (CFD) to neutronics simulations often incurs significant computational costs due to the complexity of surface intersection calculations in Monte Carlo codes. This paper presents an accelerated multiphysics coupling method for neutronics that utilizes hierarchical agglomerative clustering to map complex material property distributions to a neutronics model. Implemented within the fusion reactor design and assessment (FREDA) framework, the method leverages existing Python packages to automate the creation of clustered geometries for OpenMC. The approach is demonstrated on a sector model of an ARC-class tokamak with an immersion molten salt blanket, and an simple geometry with varying isotopic concentrations. Results show that the clustering method significantly reduces computational burden without compromising fidelity, providing a foundation for agile iteration of neutronics simulations involving multiple coupled material properties.

Bae, Jin Whan [ORNL] (ORCID:0000000326548907)↗

MONACO: accurate biological network alignment through optimal neighborhood matching between focal nodes

Motivation: Alignment of protein–protein interaction networks can be used for the unsupervised prediction of functional modules, such as protein complexes and signaling pathways, that are conserved across different species. To date, various algorithms have been proposed for biological network alignment, many of which attempt to incorporate topological similarity between the networks into the alignment process with the goal of constructing accurate and biologically meaningful alignments. Especially, random walk models have been shown to be effective for quantifying the global topological relatedness between nodes that belong to different networks by diffusing node-level similarity along the interaction edges. However, these schemes are not ideal for capturing the local topological similarity between nodes. Results: Here, we propose MONACO, a novel and versatile network alignment algorithm that finds highly accurate pairwise and multiple network alignments through the iterative optimal matching of ‘local’ neighborhoods around focal nodes. Extensive performance assessment based on real networks as well as synthetic networks, for which the ground truth is known, demonstrates that MONACO clearly and consistently outperforms all other state-of-the-art network alignment algorithms that we have tested, in terms of accuracy, coherence and topological quality of the aligned network regions. Furthermore, despite the sharply enhanced alignment accuracy, MONACO remains computationally efficient and it scales well with increasing size and number of networks.

97 MATHEMATICS AND COMPUTING↗

Safeguards-Informed Hybrid Imagery Dataset [Poster]

Deep Learning computer vision models require many thousands of properly labelled images for training, which is especially challenging for safeguards and nonproliferation, given that safeguards-relevant images are typically rare due to the sensitivity and limited availability of the technologies. Creating relevant images through real-world staging is costly and limiting in scope. Expert-labeling is expensive, time consuming, and error prone. We aim to develop a data set of both realworld and synthetic images that are relevant to the nuclear safeguards domain that can be used to support multiple data science research questions. In the process of developing this data, we aim to develop a novel workflow to validate synthetic images using machine learning explainability methods, testing among multiple computer vision algorithms, and iterative synthetic data rendering. We will deliver one million images – both real-world and synthetically rendered – of two types uranium storage and transportation containers with labelled ground truth and associated adversarial examples.

97 MATHEMATICS AND COMPUTING↗

Progress Towards A Measurement of Neutrino Induced Charged Current Neutral Pion Production in the MicroBooNE Experiment

An analysis of MicroBooNE data with a signal of one muon, one neutral pion, and no charged pions is presented. Studying neutral pion production in the MicroBooNE detector provides an opportunity to better understand neutrino-argon interactions, and is crucial for future accelerator-based neutrino oscillation experiments. This analysis presents the progress towards the first measurement of the differential cross section for charged current (CC) $π^0$ production in neutrino-argon interactions. Using a dataset corresponding to about 7 × 10 20 protons on target (POT), we present an analysis which aims to measure the single differential cross sections as a function of the $π^0$ kinematic variables such as momentum and scattering angle. The Wiener-SVD technique for unfolding the measurement is presented and demonstrated using multiple generator predictions. A future iteration of this analysis will compare an unfolded data measurement to these models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Real-space solution to the electronic structure problem for nearly a million electrons

We report a Kohn–Sham density functional theory calculation of a system with more than 200 000 atoms and 800 000 electrons using a real-space high-order finite-difference method to investigate the electronic structure of large spherical silicon nanoclusters. Our system of choice was a 20 nm large spherical nanocluster with 202 617 silicon atoms and 13 836 hydrogen atoms used to passivate the dangling surface bonds. To speed up the convergence of the eigenspace, we utilized Chebyshev-filtered subspace iteration, and for sparse matrix–vector multiplications, we used blockwise Hilbert space-filling curves, implemented in the PARSEC code. For this calculation, we also replaced our orthonormalization + Rayleigh–Ritz step with a generalized eigenvalue problem step. We utilized all of the 8192 nodes (458 752 processors) on the Frontera machine at the Texas Advanced Computing Center. We achieved two Chebyshev-filtered subspace iterations, yielding a good approximation of the electronic density of states. Our work pushes the limits on the capabilities of the current electronic structure solvers to nearly 106 electrons and demonstrates the potential of the real-space approach to efficiently parallelize large calculations on modern high-performance computing platforms.

Chemistry↗

Subcritical Multiplication with a Fixed Source

In a subcritical, multiplying medium, the system multiplication describes the expected total number of neutrons created by a single source neutron. Subcriticality plays a large role in criticality safety and thus it is vital for the subcritical multiplication factor be accurate, especially as a system approaches criticality. This work examines the accuracy of calculating the system multiplication using the MCNP6.2 ® k-eigenvalue power iteration (KCODE) method when a fixed-point source is present in a multiplying medium, for near critical systems. This work compares the standard approach for calculating system multiplication, using the fixed-source calculational approach, to a new, single k-eigenvalue power iteration approach that incorporates a fixed-source component and a fission-source component into a single calculation. For the remainder of this paper, some theoretical background and numerical results for an approximate k eigenvalue approach, an accurate fixed-source approach and a new and more accurate k-eigenvalue approach to computing system multiplication are provided.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The interaction of the ITER first wall with magnetic perturbations

Mitigation of the multiple risks associated with disruptions and runaway electrons in tokamaks involves competing demands. Success requires that each risk be understood sufficiently that appropriate compromises can be made. Here the focus is on the interaction of short timescale magnetic-perturbations with the structure in ITER that is closest to the plasma, blanket modules covered by separated beryllium tiles. The effect of this tiled surface on the perturbations and on the forces on structures is subtle. Indeterminacy can be introduced by tile-to-tile shorting. A determinate subtlety is introduced because electrically separated tiles can act as a conducting surface for magnetic perturbations that have a normal component to the surface. A practical method for including this determinate subtlety into plasma simulations is developed. The shorter the timescales and the greater the localization, particularly in the toroidal direction, the more important the magnetic effects of the tiles become.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Fox Trails

1. This software utilizes python pandas to pull data from P6 databases or XER files. The software transforms the datasets into multiple main tables by joining, filtering, iteratively flattening hierarchical structured data, and pivoting datasets to give simple flat output tables. The activity table includes all of the information related to an activity including activity codes, global, EPS, and project codes, UDFs, and WBS information as separate columns. This includes the code id, code value and sequence number for all levels in hierarchical codes. The resource table is similar to the activity table and includes all of the information related to resources on activities including UPFs and resource codes. The resource time phased table takes the resource information and time phases it for the budget, forecast, late, and actual dates/units/costs that closely matches P6's user interface's values as it implements the resource curve and calendars. The wbs table contains the WBS structure broken out by levels and includes UDFs, codes, and notebook topics. The final P6 data table is the relationships table which simply contains the relationships. 2. When a user updates the tool with data (via giving it P6 project names with database username/password information or XER files) the system creates the data in #1, then creates a networkx graph with the activity data imbedded in the node data and the relationships added as edges. Each edge also has it's float calculated (working time distance between the predecessor and successor) and attached to the edge. Activities are also tagged as a potential start of a path based on their constraints, constraint dates, remaining start date, and activity status. When a user enters an activity ID into the UI, it runs a shortest path calculation on the network graph between each node tagged as potential start to the entered activity id based on the float tagged on the edge. Each path returned by the algorithm contains all of the nodes on the path in order, as well as the total float of the edges that make the path. This data is then collected and returned to the user in the form of a gantt chart with groupings for each path that includes the total float for each group. 3. Similar to 2, if the user passes through a reference dataset each activity set in the path is checked to see if it had a path in the reference dataset, if that path was the primary path between the start and end activities, and what has changed regarding logic and durations. These changes are color coded and summarized before sent to the user to be displayed by the UI for simple discovery. 4. Utilizing the data from #1, the user can submit desired grouping code(s) and filters to the system. The system will then pull the activities, resources, and relationships and create a gantt chart based on the groupings sent and filtered based on the filters sent. 5. The system will produce a gantt chart in a similar method to #4, but allows interactivity with the data. As the user interacts with the gantt chart, the software captures the changes and stores it with the user making the change so that project controls and implement those changes in P6.

Fox, Ben↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

TETRIS-ADAPT-VQE: An adaptive algorithm that yields shallower, denser circuit Ansätze

Adaptive quantum variational algorithms are particularly promising for simulating strongly correlated systems on near-term quantum hardware, but they are not yet viable due, in large part, to the severe coherence time limitations on current devices. In this paper, we introduce an algorithm called TETRIS-ADAPT-VQE (tiling efficient trial circuits with rotations implemented simultaneously adaptive derivative-assembled problem-tailored variational quantum eigensolver), which iteratively builds up variational a few operators at a time in a way dictated by the problem being simulated. This algorithm is a modified version of the ADAPT-VQE algorithm, in which the one-operator-at-a-time rule is lifted to allow for the addition of multiple operators with disjoint supports in each iteration. TETRIS-ADAPT-VQE results in denser but significantly shallower circuits, without increasing the number of controlled- gates or variational parameters. Its advantage over the original algorithm in terms of circuit depths increases with the system size. Moreover, the expensive step of measuring the energy gradient with respect to each candidate unitary at each iteration is performed only a fraction of the time compared with ADAPT-VQE. These improvements bring us closer to the goal of demonstrating a practical quantum advantage on quantum hardware. Published by the American Physical Society 2024

Anastasiou, Panagiotis G. (ORCID:0000000256601791)↗

HyKKT: a hybrid direct-iterative method for solving KKT linear systems

Here, we propose a solution strategy for the large indefinite linear systems arising in interior methods for nonlinear optimization. The method is suitable for implementation on hardware accelerators such as graphical processing units (GPUs). The current gold standard for sparse indefinite systems is the LBLT factorization where L is a lower triangular matrix and B is 1×1 or 2×2 block diagonal. However, this requires pivoting, which substantially increases communication cost and degrades performance on GPUs. Our approach solves a large indefinite system by solving multiple smaller positive definite systems, using an iterative solver on the Schur complement and an inner direct solve (via Cholesky factorization) within each iteration. Cholesky is stable without pivoting, thereby reducing communication and allowing reuse of the symbolic factorization. We demonstrate the practicality of our approach on large optimal power flow problems and show that it can efficiently utilize GPUs and outperform LBL T factorization of the full system.

97 MATHEMATICS AND COMPUTING↗

Response Surface Methodology As a New Approach for Finding Optimal MALDI Matrix Spraying Parameters for Mass Spectrometry Imaging

Automated spraying devices have become ubiquitous in laboratories employing matrixassisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI), in part because they permit control of a number of matrix application parameters that can easily be reproduced for intra- and interlaboratory studies. Determining the optimal parameters for MALDI matrix application, such as temperature, flow rate, spraying velocity, number of spraying cycles, and solvent composition for matrix application, is critical for obtaining high-quality MALDI-MSI data. However, there are no established approaches for optimizing these multiple parameters simultaneously. Instead optimization is performed iteratively (i.e., one parameter at a time), which is time consuming and can lead to overall non-optimal settings. In this report, we demonstrate the use a novel experimental design and the response surface methodology to optimize five parameters of MALDI matrix application using a robotic sprayer. Thirtytwo combinations of MALDI matrix spraying conditions were tested, which allowed us to elucidate relationships between each of the application parameters as determined by MALDI-MS (specifically, using a 15 Tesla Fourier transform ion cyclotron resonance mass spectrometer). As such, we were able to determine the optimal automated spraying parameters that minimized signal de-localization and enabled high MALDI sensitivity. We envision this optimization strategy can be utilized for other MALDI-MSI applications, tissue types, and matrix application approaches.

60 APPLIED LIFE SCIENCES↗

The Synthetic Biology Open Language (SBOL) Version 3: Simplified Data Exchange for Bioengineering

The Synthetic Biology Open Language (SBOL) is a community-developed data standard that allows knowledge about biological designs to be captured using a machine-tractable, ontology-backed representation that is built using Semantic Web technologies. While early versions of SBOL focused only on the description of DNA-based components and their sub-components, SBOL can now be used to represent knowledge across multiple scales and throughout the entire synthetic biology workflow, from the specification of a single molecule or DNA fragment through to multicellular systems containing multiple interacting genetic circuits. The third major iteration of the SBOL standard, SBOL3, is an effort to streamline and simplify the underlying data model with a focus on real-world applications, based on experience from the deployment of SBOL in a variety of scientific and industrial settings. Here, we introduce the SBOL3 specification both in comparison to previous versions of SBOL and through practical examples of its use.

59 BASIC BIOLOGICAL SCIENCES↗