Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multiple iterations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Revisiting Source Convergence Diagnostics in the KENO Monte Carlo Neutron Transport Codes [Abstract]

Monte Carlo criticality transport codes, which rely on the power iteration procedure, are a fundamental tool for nuclear criticality safety practitioners in assessing the neutron multiplication factor (k eff ) for problems involving fissile material. In these calculations, ensuring the convergence of both the fission source distributions and the k eff estimate for accurate results is crucial. However, a converged k eff estimate does not necessarily mean the fission source distribution is also converged because the fission source and flux distribution may continue to evolve even after k eff convergence. Therefore, most Monte Carlo transport criticality codes now offer various diagnostic tests to assess fission source convergence in addition to the k eff convergence by analyzing the trends of these quantities over multiple generations.

AZURE↗

Metabolic flux optimization of iterative pathways through orthogonal gene expression control: Application to the β-oxidation reversal

Balancing relative expression of pathway genes to minimize flux bottlenecks and metabolic burden is one of the key challenges in metabolic engineering. This is especially relevant for iterative pathways, such as reverse β-oxidation (rBOX) pathway, which require control of flux partition at multiple nodes to achieve efficient synthesis of target products. Here, we develop a plasmid-based inducible system for orthogonal control of gene expression (referred to as the TriO system) and demonstrate its utility in the rBOX pathway. Leveraging effortless construction of TriO vectors in a plug-and-play manner, we simultaneously explored the solution space for enzyme choice and relative expression levels. Remarkably, varying individual expression levels led to substantial change in product specificity ranging from no production to optimal performance of about 90% of the theoretical yield of the desired products. We obtained titers of 6.3 g/L butyrate, 2.2 g/L butanol and 4.0 g/L hexanoate from glycerol in E. coli, which exceed the best titers previously reported using equivalent enzyme combinations. Since a similar system behavior was observed with alternative termination routes and higher-order iterations, we envision our approach to be broadly applicable to other iterative pathways besides the rBOX. Here, considering that high throughput, automated strain construction using combinatorial promoter and RBS libraries remain out of reach for many researchers, especially in academia, tools like the TriO system could democratize the testing and evaluation of pathway designs by reducing cost, time and infrastructure requirements.

59 BASIC BIOLOGICAL SCIENCES↗

MULTI-LEADER: MULTI-source LEarning-Accelerated Design of high-Efficiency multi-stage compRessor (Final Technical Report)

The objective of MULTI-LEADER is to cut design costs by 80% while generating more energy-efficient designs of multi-stage compressors by developing and implementing novel machine learning (ML) techniques, which enable faster and fewer design iterations, improved solver performance, and concurrent multi-disciplinary design. Current industrial practices for the design of multi-stage compressors involve simulation-based design optimization with successive levels of model fidelity, iteratively evaluated between distinct disciplines, one stage at a time to tackle the high dimensional design variations. This project addresses these key design challenges: (1) concurrent optimization of multiple stages under many non-linear constraints; (2) multitude of evaluation of high-fidelity and expensive solvers and their gradients during optimization convergence in high-dimensional design; (3) multi-disciplinary design to maximize aerodynamic performance while guaranteeing structural integrity and additive manufacturability; (4) utilization of multiple fidelity of solvers with disparate parameterization and modeling assumptions. MULTI-LEADER achieved more than 5x speed up in detailed design of more energy-efficient compressors via these machine learning (ML) innovations: (i) rapid design surrogates by multi-source learning from diverse fidelities across multiple disciplines, (ii) physics-constrained data-augmented modeling for improved empiricism, (iii) generative manifold embedding for high dimensional concurrent design without gradient information; (iv) budget-constrained fidelity-adaptive sampling towards fewer design iterations.

33 ADVANCED PROPULSION SYSTEMS↗

Single and multiple scattering contributions to circumsolar radiation

The contributions to the angular distribution of the almucantar radiance in the forward direction due to multiple scattering are compared to those due to single scattering. The contributions have been calculated by a computer code employing the Gauss-Seidel iterative approach to the solution of the radiative transfer equation for a plane parallel atmosphere composed of air molecules, aerosol particles, and ozone. The code is similar to that of Dave (1972) except in the construction of the source matrix. In the near-forward direction the multiple scattering contributions are significant for optical depths of the order of 0.4. The shape of the angular distribution of almucantar radiance to 10 degrees is less sensitive to multiple scattering.

Box, M. A.↗

MS2Planner: improved fragmentation spectra coverage in untargeted mass spectrometry by iterative optimized data acquisition

Motivation: Untargeted mass spectrometry experiments enable the profiling of metabolites in complex biological samples. The collected fragmentation spectra are the metabolite’s fingerprints that are used for molecule identification and discovery. Two main mass spectrometry strategies exist for the collection of fragmentation spectra: data-dependent acquisition (DDA) and data-independent acquisition (DIA). In the DIA strategy, all the metabolites ions in predefined mass-to-charge ratio ranges are co-isolated and co-fragmented, resulting in multiplexed fragmentation spectra that are challenging to annotate. In contrast, in the DDA strategy, fragmentation spectra are dynamically and specifically collected for the most abundant ions observed, causing redundancy and sub-optimal fragmentation spectra collection. Yet, DDA results in less multiplexed fragmentation spectra that can be readily annotated. Results: We introduce the MS2Planner workflow, an Iterative Optimized Data Acquisition strategy that optimizes the number of high-quality fragmentation spectra over multiple experimental acquisitions using topological sorting. Our results showed that MS2Planner increases the annotation rate by 38.6% and is 62.5% more sensitive and 9.4% more specific compared to DDA. Availability and implementation MS2Planner code is available at https://github.com/mohimanilab/MS2Planner. The generation of the inclusion list from MS2Planner was performed with python scripts available at https://github.com/lfnothias/IODA_MS.

47 OTHER INSTRUMENTATION↗

TunIO: An AI-powered Framework for Optimizing HPC I/O

I/O operations are a known performance bottleneck of HPC applications. To achieve good performance, users often employ an iterative multistage tuning process to find an optimal I/O stack configuration. However, an I/O stack contains multiple layers, such as high-level I/O libraries, I/O middleware, and parallel file systems, and each layer has many parameters. These parameters and layers are entangled and influenced by each other. The tuning process is time-consuming and complex. In this work, we present TunIO, an AI-powered I/O tuning framework that implements several techniques to balance the tuning cost and performance gain, including tuning the high-impact parameters first. Furthermore, TunIO analyzes the application source code to extract its I/O kernel while retaining all statements necessary to perform I/O. It utilizes a smart selection of high-impact configuration parameters of the given tuning objective. Finally, it uses a novel Reinforcement Learning (RL)-driven early stopping mechanism to balance the cost and performance gain. Experimental results show that TunIO leads to a reduction of up to ≈73% in tuning time while achieving the same performance gain when compared to H5Tuner. It achieves a significant performance gain/cost of 208.4 MBps/min (I/O bandwidth for each minute spent in tuning) over existing approaches under our testing.

Rajesh, Neeraj↗

Model-Based Reconstruction for Multi-Frequency Collimated Beam Ultrasound Systems

Collimated beam ultrasound systems are a technology for imaging inside multi-layered structures such as geothermal wells. These systems work by using a collimated narrow-band ultrasound transmitter that can penetrate through multiple layers of heterogeneous material. A series of measurements can then be made at multiple transmit frequencies. However, commonly used reconstruction algorithms such as Synthetic Aperture Focusing Technique (SAFT) tend to produce poor quality reconstructions for these systems both because they do not model collimated beam systems and they do not jointly reconstruct the multiple frequencies. Here, in this article, we propose a multi-frequency ultrasound model-based iterative reconstruction (UMBIR) algorithm designed for multi-frequency collimated beam ultrasound systems. The combined system targets reflective imaging of heterogeneous, multi-layered structures. For each transmitted frequency band, we introduce a physics-based forward model to accurately account for the propagation of the collimated narrow-band ultrasonic beam through the multi-layered media. We then show how the joint multi-frequency UMBIR reconstruction can be computed by modeling the direct arrival signals, detector noise, and incorporating a spatially varying image prior. Results using both simulated and experimental data indicate that multi-frequency UMBIR reconstruction yields much higher reconstruction quality than either single frequency UMBIR or SAFT.

47 OTHER INSTRUMENTATION↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Inverse Aerodynamic Design of Gas Turbine Blades using Probabilistic Machine Learning

Abstract One of the critical components in Industrial Gas Turbines (IGT) is the turbine blade. Design of turbine blades needs to consider multiple aspects like aerodynamic efficiency, durability, safety and manufacturing, which make the design process sequential and iterative. The sequential nature of these iterations forces a long design cycle time, ranging from several months to years. Due to the reactionary nature of these iterations, little effort has been made to accumulate data in a manner that allows for deep exploration and understanding of the total design space. This is exemplified in the process of designing the individual components of the IGT resulting in a potential unrealized efficiency. To overcome the aforementioned challenges, we demonstrate a probabilistic inverse design machine learning framework, namely PMI (PMI), to carry out an explicit inverse design. PMI calculates the design explicitly without costly iteration and overcomes the challenges associated with ill-posed inverse problems. In this work the framework will be demonstrated on inverse aerodynamic design of three-dimensional turbine blades.

Engineering↗

RxnRover/amlro

AMLRO (Active Machine Learning Reaction Optimizer) is an open-source framework designed to accelerate chemical reaction optimization using active learning with classical machine learning regression models. AMLRO integrates space-filling sampling strategies (e.g., Sobol and Latin Hypercube sampling) with iterative model training, prediction, and experiment selection to efficiently navigate complex reaction spaces. The platform supports multiple regression models, flexible multi-objective definitions, and user-defined parameter bounds, enabling data-efficient optimization from small initial datasets. AMLRO is designed for ease of use by experimentalists and can operate as a standalone decision-support tool or be integrated into closed-loop automated experimentation workflows.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

Understanding performance variability in standard and pipelined parallel Krylov solvers

In this work, we collect data from runs of Krylov subspace methods and pipelined Krylov algorithms in an effort to understand and model the impact of machine noise and other sources of variability on performance. We find large variability of Krylov iterations between compute nodes for standard methods that is reduced in pipelined algorithms, directly supporting conjecture, as well as large variation between statistical distributions of runtimes across iterations. Based on these results, we improve upon a previously introduced nondeterministic performance model by allowing iterations to fluctuate over time. We present our data from runs of various Krylov algorithms across multiple platforms as well as our updated non-stationary model that provides good agreement with observations. We also suggest how it can be used as a predictive tool.

97 MATHEMATICS AND COMPUTING↗

A parallel iterative solution method for systems of nonlinear hyperbolic equations

An iterative algorithm suitable for the solution of a system of nonlinear hyperbolic partial differentiation equations in multiple dimensions is discussed. Current numerical methods for systems of nonlinear PDEs have limited parallelism due to strong coupling between the equations. This method decouples the PDEs by linearizing the convention coefficient for a space-time domain. This provides large grain parallelism. The linearization also allows the treatment of some terms in the equations as source terms, providing more freedom to choose from a wider variety of numerical methods. Smaller grain parallelism may be exploited within the solves for each equation. Thus, the method has potential for parallelism at several levels.

Scroggs, Jeffrey S.↗

Domain Decomposition Algorithms for First-Order System Least Squares Methods

Least squares methods based on first-order systems have been recently proposed and analyzed for second-order elliptic equations and systems. They produce symmetric and positive definite discrete systems by using standard finite element spaces, which are not required to satisfy the inf-sup condition. In this paper, several domain decomposition algorithms for these first-order least squares methods are studied. Some representative overlapping and substructuring algorithms are considered in their additive and multiplicative variants. The theoretical and numerical results obtained show that the classical convergence bounds (on the iteration operator) for standard Galerkin discretizations are also valid for least squares methods.

Pavarino, Luca F.↗

Parallel Preconditioning for CFD Problems on the CM-5

Up to today, preconditioning methods on massively parallel systems have faced a major difficulty. The most successful preconditioning methods in terms of accelerating the convergence of the iterative solver such as incomplete LU factorizations are notoriously difficult to implement on parallel machines for two reasons: (1) the actual computation of the preconditioner is not very floating-point intensive, but requires a large amount of unstructured communication, and (2) the application of the preconditioning matrix in the iteration phase (i.e. triangular solves) are difficult to parallelize because of the recursive nature of the computation. Here we present a new approach to preconditioning for very large, sparse, unsymmetric, linear systems, which avoids both difficulties. We explicitly compute an approximate inverse to our original matrix. This new preconditioning matrix can be applied most efficiently for iterative methods on massively parallel machines, since the preconditioning phase involves only a matrix-vector multiplication, with possibly a dense matrix. Furthermore the actual computation of the preconditioning matrix has natural parallelism. For a problem of size n, the preconditioning matrix can be computed by solving n independent small least squares problems. The algorithm and its implementation on the Connection Machine CM-5 are discussed in detail and supported by extensive timings obtained from real problem data.

Simon, Horst D.↗

Real-Time Parameter Estimation for Flexible Aircraft

A method for estimating aeroelastic stability and control derivatives for flexible aircraft is developed and demonstrated using flight test data for the X-56A subscale demonstrator. The method uses the equation-error approach with frequency-domain data, and can be applied post-flight or in real time during flight. The non-dimensional aeroelastic forces and moments and the explanatory variables (including generalized displacement, rate, and acceleration states for the vibration modes) are estimated using a finite element model and onboard sensor measurements in both a least squares and Kalman filtering framework. The data are then transformed into the frequency domain for parameter estimation using equation error. This method can result in a more efficient analysis than with other iterative methods, and can leverage existing statistical tools for model structure determination, data collinearity detection, combining multiple maneuvers or prior information, and others to improve model quality.

Grauer, Jared A.↗

A hypermatrix formulation for subspace iteration

The computational efficiency of subspace iteration is addressed relative to the data structures adopted for the very large and generally sparse coefficient matrices. The frequent triangulations and matrix multiplications demand that access to the terms in the coefficient matrices be unbiased. Reliance on virtual memory (paging) operating systems with no special considerations for localized data access is not adequate. Specific data structures must be designed that accommodate the needs of the numerical algorithm yet eliminate unnecessary paging. An implementation of the subspace iteration method using hypermatrix data structures is presented. Use of hypermatrices is shown to provide unbiased and localized data access. The various modifications to the conventional formulation are described and an example problem illustrates the potential benefits of the hypermatrix formulation. Possibilities for adapting hypermatrix data structures to new supercomputer architectures are discussed.

Schmidt, Richard J.↗

Registration of video sequences from multiple sensors

In this paper, we describe an approach for registration of video sequences from a suite of multiple sensors including television, infrared and radar. Video sequences generated by these sensors may contain abrupt changes in local contrast and inconsistent image features, which pose additional difficulties for registration. Our approach to registration addresses the difficulties caused by using multiple sensors. We use a representation for registration that is invariant to local contrast changes, followed by smoothing of the resulting error measure used for registration, for robust estimation of registration parameters. We use an iterative procedure to reduce the effect of inconsistent features. Finally, we describe a method that uses same-sensor registration to aide in performing registration of sequences of video frames across multiple sensors.

Sharma, Ravi K.↗

Novel SOLPS-ITER simulations of X-point target and snowflake divertors

Abstract The design and understanding of alternative divertor configurations may be crucial for achieving acceptable steady-state heat and particle material loads for magnetic confinement fusion reactors. Multiple X-point alternative divertor geometries such as snowflakes and X-point targets have great potential in reducing power loads, but have not yet been simulated widely in codes with kinetic neutrals. This paper discusses recent changes made to the SOLPS-ITER code to allow for the simulation of X-point target and low-field side snowflake divertor geometries. Snowflake simulations using this method are presented, in addition to the first SOLPS-ITER simulation of the X-point target. Analysis of these results show reasonable consistency with the simple modelling and theoretical predictions, supporting the validity of the methodology implemented.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗