Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “experimental algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Explaining Missing Data in Graphs: A Constraint-based Approach

Abstract: This paper introduces a constraint-based approach to clarify missing values in graphs. Our method capitalizes on a set S of graph data constraints. An explanation is a sequence of operational enforcement of S towards the recovery of interested yet missing data (e.g., attribute values, edges). We show that constraint-based approach helps us to understand not only why a value is missing, but also how to recover the missing value. We study S-explanation problem, which is to compute the optimal explanations with guarantees on the informativeness and conciseness. We show the problem is in ?P^2 for established graph data constraints such as graph keys and graph association rules. We develop an efficient bidirectional algorithm to compute optimal explanations, without enforcing S on the entire graph. We also show our algorithm can be easily extended to support graph refinement within limited time, and to explain missing answers. Using real-world graphs, we experimentally verify the effectiveness and efficiency of our algorithms.

Data Analytics↗

Direction-optimizing Label Propagation Framework for Structure Detection in Graphs: Design, Implementation, and Experimental Analysis

Label Propagation is not only a well-known machine learning algorithm for classification but also an effective method for discovering communities and connected components in networks. We propose a new Direction-optimizing Label Propagation Algorithm (DOLPA) framework that enhances the performance of the standard Label Propagation Algorithm (LPA), increases its scalability, and extends its versatility and application scope. As a central feature, the DOLPA framework relies on the use of frontiers and alternates between label push and label pull operations to attain high performance. It is formulated in such a way that the same basic algorithm can be used for finding communities or connected components in graphs by only changing the objective function used. Additionally, DOLPA has parameters for tuning the processing order of vertices in a graph to reduce the number of edges visited and improve the quality of solution obtained. We present the design and implementation of the enhanced algorithm as well as our shared-memory parallelization of it using OpenMP. We also present an extensive experimental evaluation of our implementations using the LFR benchmark and real-world networks drawn from various domains. Compared with an implementation of LPA for community detection available in a widely used network analysis software, we achieve at most five times the F-Score while maintaining similar runtime for graphs with overlapping communities. We also compare DOLPA against an implementation of the Louvain method for community detection using the same LFR-graphs and show that DOLPA achieves about three times the F-Score at just 10% of the runtime. For connected component decomposition, our algorithm achieves orders of magnitude speedups over the basic LP-based algorithm on large-diameter graphs, up to 13.2× speedup over the Shiloach-Vishkin algorithm, and up to 1.6× speedup over Afforest on an Intel Xeon processor using 40 threads.

97 MATHEMATICS AND COMPUTING↗

Bayesian optimization algorithms for accelerator physics

Accelerator physics relies on numerical algorithms to solve optimization problems in online accelerator control and tasks such as experimental design and model calibration in simulations. The effectiveness of optimization algorithms in discovering ideal solutions for complex challenges with limited resources often determines the problem complexity these methods can address. The accelerator physics community has recognized the advantages of Bayesian optimization algorithms, which leverage statistical surrogate models of objective functions to effectively address complex optimization challenges, especially in the presence of noise during accelerator operation and in resource-intensive physics simulations. In this review article, we offer a conceptual overview of applying Bayesian optimization techniques toward solving optimization problems in accelerator physics. We begin by providing a straightforward explanation of the essential components that make up Bayesian optimization techniques. We then give an overview of current and previous work applying and modifying these techniques to solve accelerator physics challenges. Finally, we explore practical implementation strategies for Bayesian optimization algorithms to maximize their performance, enabling users to effectively address complex optimization challenges in real-time beam control and accelerator design. Published by the American Physical Society 2024

43 PARTICLE ACCELERATORS↗

Complex Oxides for Brain–Inspired Computing: A Review

The fields of brain-inspired computing, robotics, and, more broadly, artificial intelligence (AI) seek to implement knowledge gleaned from the natural world into human-designed electronics and machines. In this review, the opportunities presented by complex oxides, a class of electronic ceramic materials whose properties can be elegantly tuned by doping, electron interactions, and a variety of external stimuli near room temperature, are discussed. The review begins with a discussion of natural intelligence at the elementary level in the nervous system, followed by collective intelligence and learning at the animal colony level mediated by social interactions. An important aspect highlighted is the vast spatial and temporal scales involved in learning and memory. The focus then turns to collective phenomena, such as metal-to-insulator transitions (MITs), ferroelectricity, and related examples, to highlight recent demonstrations of artificial neurons, synapses, and circuits and their learning. First-principles theoretical treatments of the electronic structure, and in situ synchrotron spectroscopy of operating devices are then discussed. The implementation of the experimental characteristics into neural networks and algorithm design is then revewed. Finally, outstanding materials challenges that require a microscopic understanding of the physical mechanisms, which will be essential for advancing the frontiers of neuromorphic computing, are highlighted.

36 MATERIALS SCIENCE↗

Estimating Sediment Settling Velocities from a Theoretically Guided Data-Driven Approach

Sediment settling velocities are commonly estimated from analytical or process-based approaches. These approaches have theoretical constraints due to the incompletely resolved settling physics. A parametric data-driven approach was recently proposed without theoretical constraints, but it is limited by its mathematical assumptions. To overcome these limitations, here we apply a machine learning algorithm to an aggregated sediment settling experimental database and develops a nonparametric data-driven model to estimate the noncohesive sediment settling velocity in water. A cross-comparison against five process-based equations and a parametric data-driven equation demonstrates the higher accuracy and better consistency of the new model in estimating sediment settling velocities under various physical regimes. The new model also shows an easily implemented self-update capability by assimilating theoretical data derived from the process-based equations. The updated model, incorporating experimental and theoretical data of sediment settling processes, further improves the accuracy and reduces the uncertainty in estimating sediment settling velocities. This study demonstrates the capability of machine learning in sediment transport study and illustrates an alternative framework for other hydraulic engineering challenges.

42 ENGINEERING↗

Machine learning methods for probabilistic locked-mode predictors in tokamak plasmas

A rotating tokamak plasma can interact resonantly with the external helical magnetic perturbations, also known as error fields. This can lead to locking and then to disruptions. We leverage machine learning (ML) methods to predict the locking events. We use a coupled third-order nonlinear ordinary differential equation model to represent the interaction of the magnetic perturbation and the plasma rotation with the error field. This model is sufficient to describe qualitatively the locking and unlocking bifurcations. Here, we explore using ML algorithms with the simulation data and experimental data, focusing on the methods that can be used with sparse datasets. These methods lead to the possibility of the avoidance of locking in real-time operations. We describe the operational space in terms of two control parameters: the magnitude of the error field and the rotation frequency associated with the momentum source that maintains the plasma rotation. The outcomes are quan- tified by order parameters that completely characterize the state, whether locked or unlocked. We use unsupervised ML methods to classify locked/unlocked states and note the usefulness of a certain normalization of the order parameters. Three supervised ML classifiers are used in suite to estimate the probability of locking in the region of control parameter space with hysteresis, i.e., the set of control parameters for which both locked and unlocked states can exist. The results show that a neural network gives the best estimate of the locking probability. An analogy of the present locking model with the van der Waals equation of state is also provided.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Efficient numerical algorithm for multi-level ionization of high-atomic-number gases

An efficient numerical algorithm for laser driven multi-level ionization of high-atomic-number gases is proposed and implemented in an electromagnetic particle-in-cell code SPACE. The algorithm is based on analytical solutions to the system of differential equations describing ionization evolution. Using analytical solutions resolves the multiscale issue of ionization due to different characteristic time scales of ionization processes and the main code time step. Algorithm efficiency and memory requirements are significantly improved by using a locally reduced system of differential equations. The algorithm also assigns proper orbital quantum numbers and their projections to ionization states. The algorithm is verified and validated using experimental data.

Cheng, A. (ORCID:000000021945282X)↗

Accelerating the discovery of novel magnetic materials using machine learning–guided adaptive feedback

Magnetic materials are essential for energy generation and information devices, and they play an important role in advanced technologies and green energy economies. Currently, the most widely used magnets contain rare earth (RE) elements. An outstanding challenge of notable scientific interest is the discovery and synthesis of novel magnetic materials without RE elements that meet the performance and cost goals for advanced electromagnetic devices. Here, we report our discovery and synthesis of an RE-free magnetic compound, Fe 3 CoB 2 , through an efficient feedback framework by integrating machine learning (ML), an adaptive genetic algorithm, first-principles calculations, and experimental synthesis. Magnetic measurements show that Fe 3 CoB 2 exhibits a high magnetic anisotropy ( K 1 = 1.2 MJ/m 3 ) and saturation magnetic polarization ( J s = 1.39 T), which is suitable for RE-free permanent-magnet applications. Our ML-guided approach presents a promising paradigm for efficient materials design and discovery and can also be applied to the search for other functional materials.

36 MATERIALS SCIENCE↗

Progress in modelling fast-ion D-alpha spectra and neutral particle analyzer fluxes using FIDASIM

FIDASIM is a code that models signals produced by charge-exchange reactions between neutrals and ions (both fast and thermal) in magnetically confined plasmas. With the ion distribution function as input, the code predicts the efflux to a neutral particle analyzer diagnostic and the photon radiance of Balmer-alpha light to a fast-ion D α diagnostic, in addition to many other related quantities. A new, parallelized version of the Monte Carlo code FIDASIM has been developed in Fortran90 that is substantially faster than the original interactive data language version. Modified algorithms include more accurate treatments of the time dependent collisional-radiative equations that describe neutral energy levels, of the cloud of ‘halo’ neutrals that surround the injected neutral beam, and of finite Larmor radius effects. Enhanced physics capabilities include modelling ‘passive’ signals from cold edge neutrals, the ability to treat general three-dimensional magnetic confinement configurations, and calculations of diagnostic-specific weight functions that enable tomographic reconstructions of the fast-ion distribution function. Neutral beam attenuation, beam emission, and fast-ion birth profiles are also modelled. Finally, the new algorithms have been successfully validated against experimental data and new features have been tested through benchmarks between two independently developed versions of the code.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Impurity transport studies at the HSX stellarator using active and passive CVI spectroscopy

The transport of carbon impurities has been studied in the helically symmetric stellarator experiment (HSX) using active and passive charge exchange recombination spectroscopy (CHERS). For the analysis of the CHERS signals, the STRAHL impurity transport code has been re-written in the python programming language and optimized for the application in stellarators. In addition, neutral hydrogen densities both along the NBI line of sight as well as for the background plasma have been calculated using the FIDASIM code. By using the basinhopping algorithm to minimize the difference between experimental and predicted active and passive signals, significant levels of impurity diffusion are observed. In this work, comparisons with neoclassical calculations from DKES/PENTA show that the inferred levels exceed the neoclassical transport by about a factor of four in the core and more than 100 times towards the plasma edge, thus indicating a high level of anomalous transport. This observation is in agreement with experimental heat diffusivites determined from a power balance analysis which exhibits strong anomalous transport as well.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A lightweight, user-configurable detector ASIC digital architecture with on-chip data compression for MHz X-ray coherent diffraction imaging

Today, most X-ray pixel detectors used at light sources transmit raw pixel data off the detector ASIC. With the availability of more advanced ASIC technology nodes for scientific application, more digital functionalities from the computing domains (e.g., compression) can be integrated directly into a detector ASIC to increase data velocity. In this paper, we describe a lightweight, user-configurable detector ASIC digital architecture with on-chip compression which can be implemented in 130 nm technologies in a reasonable area on the ASIC periphery. In addition, we present a design to efficiently handle the variable data from the stream of parallel compressors. The architecture includes user-selectable lossy and lossless compression blocks. The impact of lossy compression algorithms is evaluated on simulated and experimental X-ray ptychography datasets. This architecture is a practical approach to increase pixel detector frame rates towards the continuous 1 MHz regime for not only coherent imaging techniques such as ptychography, but also for other diffraction techniques at X-ray light sources.

47 OTHER INSTRUMENTATION↗

CFM-ID 4.0 – a web server for accurate MS-based metabolite identification

The CFM-ID 4.0 web server (https://cfmid.wishartlab.com) is an online tool for predicting, annotating and interpreting tandem mass (MS/MS) spectra of small molecules. It is specifically designed to assist researchers pursuing studies in metabolomics, exposomics and analytical chemistry. More specifically, CFM-ID 4.0 supports the: 1) prediction of electrospray ionization quadrupole time-of-flight tandem mass spectra (ESI-QTOF-MS/MS) for small molecules over multiple collision energies (10 eV, 20 eV, and 40 eV); 2) annotation of ESI-QTOF-MS/MS spectra given the structure of the compound; and 3) identification of a small molecule that generated a given ESI-QTOF-MS/MS spectrum at one or more collision energies. The CFM-ID 4.0 web server makes use of a substantially improved MS fragmentation algorithm, a much larger database of experimental and in silico predicted MS/MS spectra and improved scoring methods to offer more accurate MS/MS spectral prediction and MS/MS-based compound identification. Compared to earlier versions of CFM-ID, this new version has an MS/MS spectral prediction performance that is ~22% better and a compound identification accuracy that is ~35% better on a standard (CASMI 2016) testing dataset. CFM-ID 4.0 also features a neutral loss function that allows users to identify similar or substituent compounds where no match can be found using CFM-ID’s regular MS/MS-to-compound identification utility. Finally, the CFM-ID 4.0 web server now offers a much more refined user interface that is easier to use, supports molecular formula identification (from MS/MS data), provides more interactively viewable data (including proposed fragment ion structures) and displays MS mirror plots for comparing predicted with observed MS/MS spectra. These improvements should make CFM-ID 4.0 much more useful to the community and should make small molecule identification much easier, faster, and more accurate.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient analysis of small-angle scattering curves for large biomolecular assemblies using Monte Carlo methods

Structure elucidation from small-angle scattering curves of large biomolecular assemblies is notoriously challenging. This is because the simulation of high-resolution features in the structure of large macromolecular assemblies, such as de novo protein assemblies, is computationally demanding when it needs to cover a broad range of length scales. Conventional methods, such as the numerical approximation to the Debye equation or the use of spherical harmonics, do not scale well as the size of the assembly increases, which limits their application to small structures (e.g. individual proteins). This work explores the effectiveness of a Monte Carlo method to simulate and fit scattering curves for large biomolecular assemblies spanning over ranges covering atomic and molecular detail (e.g. spacing and orientation of proteins in an assembly) as well as large-scale (hundreds of nanometres) features. Owing to its speed and scalability, it can be combined with a fitting algorithm to extract structural features from experimental small-angle scattering curves in biomolecular assemblies that are otherwise intractable for interpretation. This work first demonstrates the effectiveness of the tool using experimental small-angle X-ray scattering (SAXS) data from tile-like proteins that assemble into 1D tube-like macromolecular structures. Here, the diameter distribution of tubes is extracted from SAXS fits, and this is quantitatively compared with distributions from electron microscopy. SAXS data are also obtained from 2D sheet-like protein assemblies, and the proposed method is used to quantify structural features such as the separation distance between protein building blocks and the flexing of the sheet. An open-source implementation of the methodology is provided for use in a broad range of biological systems involving multi-scale scattering analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Device-Centric Firmware Malware Detection for Smart Inverters using Deep Transfer Learning

Since future power grids are inverter-dominant grids and inverters are getting smarter by incorporating remote access and seamless firmware update, it is anticipated that malware attackers will directly target smart inverters. However, malware threats targeting smart inverters have been less studied yet. This paper explores potential malware attacks targeting smart inverters and proposes a deep transfer-learning (DTL)-based malware detection framework for smart inverters. The proposed DTL method can significantly reduce development time and efforts for an artificial intelligence-based malware detection algorithm while improving detection accuracy. The experimental result shows that the proposed method achieves 98% of firmware malware detection accuracy. Furthermore, this approach will be transformative to other smart grid devices enabling seamless firmware update.

artificial intelligence↗

Accelerating Advanced Light Source Science Through Multi-Facility HPC Workflows

Synchrotron light sources support a wide array of techniques to investigate materials, often producing complex, high-volume data that challenge traditional workflows. At the Advanced Light Source (ALS), we developed infrastructure to move microtomography data over ESnet to ALCF and NERSC, where CPU- and GPU-based algorithms generate 3D reconstructed volumes of experimental samples. We employ two data movement and reconstruction models: real-time processing as data streams directly to NERSC compute nodes, and automated file transfer to NERSC and ALCF file systems. The streaming pipeline provides users with feedback in under ten seconds, while the file-based workflow produces high-quality reconstructions suitable for deeper analysis in 20-30 minutes. This infrastructure enables users to utilize HPC resources without direct access to backend systems. We plan to extend this architecture to more endstations, supporting our beamline scientists and users.

Abramov, David↗

Adaptive language model training for molecular design

Abstract The vast size of chemical space necessitates computational approaches to automate and accelerate the design of molecular sequences to guide experimental efforts for drug discovery. Genetic algorithms provide a useful framework to incrementally generate molecules by applying mutations to known chemical structures. Recently, masked language models have been applied to automate the mutation process by leveraging large compound libraries to learn commonly occurring chemical sequences (i.e., using tokenization) and predict rearrangements (i.e., using mask prediction). Here, we consider how language models can be adapted to improve molecule generation for different optimization tasks. We use two different generation strategies for comparison, fixed and adaptive. The fixed strategy uses a pre-trained model to generate mutations; the adaptive strategy trains the language model on each new generation of molecules selected for target properties during optimization. Our results show that the adaptive strategy allows the language model to more closely fit the distribution of molecules in the population. Therefore, for enhanced fitness optimization, we suggest the use of the fixed strategy during an initial phase followed by the use of the adaptive strategy. We demonstrate the impact of adaptive training by searching for molecules that optimize both heuristic metrics, drug-likeness and synthesizability, as well as predicted protein binding affinity from a surrogate model. Our results show that the adaptive strategy provides a significant improvement in fitness optimization compared to the fixed pre-trained model, empowering the application of language models to molecular design tasks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ecosystems and Networks Integrated with Genes and Molecular Assemblies (ENIGMA): Component 5: Imaging Protein Conformations, Shapes & Assemblies in Solution & Administration project (Final Scientific/Technical Report, Subcontract No. 6974584)

We set ambitious goals to examine microorganism communities and measure both their chemical input and output as a read out of specific biochemical activity. These scientific goals are driving the development of sophisticated algorithms to analyze large amounts of experimental measurements made using high throughput technologies to explain and predict how the environment influences biological function at multiple scales and how the microbial systems in tum modify the environment. By examining how bacteria communities rely on symbiotic metabolic relationships for survival and reproductive success, and how these relationships consequently affect their biochemical capabilities. This was accomplished using state-of-the-art mass spectrometry-based methods, and metabolic fingerprinting approaches with a high degree of chemical specificity and sensitivity. During this period the original effort transitioned and was consolidated into what is now ENIGMA, the efforts on technology development expanded to include more untargeted metabolomics and its application to organisms on a systems level.

59 BASIC BIOLOGICAL SCIENCES↗

Ultra-High-Speed X-ray Tomography: Bridging the Gaps to See the Unknown

Rapid, time-varying, three-dimensional physics underpin numerous engineering challenges. Often, these physics occur within opaque environments, internal to a component, severely limiting applicable diagnostics. Development of novel diagnostics is necessary to understand and predict transient three-dimensional (3D) phenomena within opaque environments. This report highlights progress in four key areas leading to advancements in high-speed X-ray radiography and tomography. The first area is enabling MHz-rate imaging of energetics at the Advanced Photon Source at Argonne National Laboratory. The second is modeling a high-flux, rotating-anode X-ray source to understand the heat loads on the anode. The third effort was to develop a novel reconstruction algorithm that is validated by ground experimental tomography data and synthetic tomography data. The fourth is the development of a novel approach to two-color X-ray imaging.

42 ENGINEERING↗