Machine Learning Overview: From Theory to Practice
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Abstract Quantum circuit learning is employed to simulate quantum field theories (QFTs). Typically, when simulating QFTs with quantum computers, significant challenges are encountered due to the technical limitations of quantum devices when implementing the Hamiltonian using Pauli spin matrices. To address this challenge, quantum circuit learning is leveraged, employing a compact configuration of qubits and low‐depth quantum circuits to predict real‐time dynamics in quantum field theories. The key advantage of this approach is that a single‐qubit measurement can accurately forecast various physical parameters, including fully‐connected operators. To demonstrate the effectiveness of this method, it is used to predict quench dynamics, chiral dynamics and jet production in a 1+1‐dimensional model of quantum electrodynamics. It is found that our predictions closely align with the results of rigorous classical calculations, exhibiting a high degree of accuracy. This hybrid quantum‐classical approach illustrates the feasibility of efficiently simulating large‐scale QFTs on cutting‐edge quantum devices.
In contrast to current intelligent systems, which must be laboriously programmed for each task they are meant to perform, instructable agents can be taught new tasks and associated knowledge. This thesis presents a general theory of learning from tutorial instruction and its use to produce an instructable agent. Tutorial instruction is a particularly powerful form of instruction, because it allows the instructor to communicate whatever kind of knowledge a student needs at whatever point it is needed. To exploit this broad flexibility, however, a tutorable agent must support a full range of interaction with its instructor to learn a full range of knowledge. Thus, unlike most machine learning tasks, which target deep learning of a single kind of knowledge from a single kind of input, tutorability requires a breadth of learning from a broad range of instructional interactions. The theory of learning from tutorial instruction presented here has two parts. First, a computational model of an intelligent agent, the problem space computational model, indicates the types of knowledge that determine an agent's performance, and thus, that should be acquirable via instruction. Second, a learning technique, called situated explanation specifies how the agent learns general knowledge from instruction. The theory is embodied by an implemented agent, Instructo-Soar, built within the Soar architecture. Instructo-Soar is able to learn hierarchies of completely new tasks, to extend task knowledge to apply in new situations, and in fact to acquire every type of knowledge it uses during task performance - control knowledge, knowledge of operators' effects, state inferences, etc. - from interactive natural language instructions. This variety of learning occurs by applying the situated explanation technique to a variety of instructional interactions involving a variety of types of instructions (commands, statements, conditionals, etc.). By taking seriously the requirements of flexible tutorial instruction, Instructo-Soar demonstrates a breadth of interaction and learning capabilities that goes beyond previous instructable systems, such as learning apprentice systems. Instructo-Soar's techniques could form the basis for future 'instructable technologies' that come equipped with basic capabilities, and can be taught by novice users to perform any number of desired tasks.
Accurate hydrodynamic modeling for laser-direct-drive (LDD) inertial-confinement-fusion (ICF) relies on precise calculations of the electron thermal conduction in all target materials. The nonlocal stopping range of electrons in ICF plasmas directly influences thermal conduction; yet, no first principles model exists for the electron mean free path in the conduction-zone regime. This work utilized time-dependent stochastic density-functional theory (TD-sDFT) to calculate the electron stopping power in deuterium-tritium (DT) plasmas at (ρ, T) conditions relevant to the conduction zone and the compressed shell in ICF. Using a combination of our TD-sDFT data and already established analytical models, we developed and trained an artificial neural network to create a global model for the nonlocal electron deposition range, λ E . We compared our machine-learning (ML) based model for λ E to the currently-used modified-Lee-More model in LDD radiation-hydrodynamic codes, such as lilac, and saw an overall decrease in the deposition range. To understand the effects of λ E on LDD ICF implosion dynamics, we implemented the ML-based model into lilac; specifically, we looked at designs consistent with a current experiment on the OMEGA laser and for a newly designed LDD-ICF target for the future OMEGA-Next facility. In both cases, we saw an overall drop in predicted ablation pressure, peak areal density, and neutron yield due to the reduced thermal conduction (smaller λ E ) in DT plasmas. Comparisons with the experiment on OMEGA are also made.
A model library containing petabytes of data is proposed by Triada, Ltd., Ann Arbor, Michigan. The library uses the newly patented N-Gram Memory Engine (Neurex), for storage, compression, and retrieval. Neurex splits data into two parts: a hierarchical network of associative memories that store 'information' from data and a permutation operator that preserves sequence. Neurex is expected to offer four advantages in mass storage systems. Neurex representations are dense, fully reversible, hence less expensive to store. Neurex becomes exponentially more stable with increasing data flow; thus its contents and the inverting algorithm may be mass produced for low cost distribution. Only a small permutation operator would be recalled from the library to recover data. Neurex may be enhanced to recall patterns using a partial pattern. Neurex nodes are measures of their pattern. Researchers might use nodes in statistical models to avoid costly sorting and counting procedures. Neurex subsumes a theory of learning and memory that the author believes extends information theory. Its first axiom is a symmetry principle: learning creates memory and memory evidences learning. The theory treats an information store that evolves from a null state to stationarity. A Neurex extracts information data without a priori knowledge; i.e., unlike neural networks, neither feedback nor training is required. The model consists of an energetically conservative field of uniformly distributed events with variable spatial and temporal scale, and an observer walking randomly through this field. A bank of band limited transducers (an 'eye'), each transducer in a bank being tuned to a sub-band, outputs signals upon registering events. Output signals are 'observed' by another transducer bank (a mid-brain), except the band limit of the second bank is narrower than the band limit of the first bank. The banks are arrayed as n 'levels' or 'time domains, td.' The banks are the hierarchical network (a cortex) and transducers are (associative) memories. A model Neurex was built and studied. Data were 50 MB to 10 GB samples of text, data base, and images: black/white, grey scale, and high resolution in several spectral bands. Memories at td, S(m(sub td)), were plotted against outputs of memories at td-1. S(m(sub td)) was Boltzman distributed, and memory frequencies exhibited self-organized criticality (SOC); i.e., 'l/f(sup beta)' after long exposures to data. Whereas output signals from level n may be encoded with B(sub output) = O(-log(2)f(sup beta)) bits, and input data encoded with B(sub input) = O((S(td)/S(td-1))(sup n)), B(sup output)/B(sub input) is much less than 1 always, the Neurex determines a canonical code for data and it is a lossless data compressor. Further tests are underway to confirm these results with more data types and larger samples.
Amid all candidates of physics beyond the Standard Model, string theory provides a unique proposal for incorporating gauge and gravitational interactions. In string theory, a four-dimensional theory that unifies quantum mechanics and gravity is obtained automatically if one posits that the additional dimensions predicted by the theory are small and curled up—a concept known as compactification. The gauge sector of the theory is specified by the topology and geometry of the extra dimensions, and the challenge is to reproduce all of the features of the Standard Model of particle physics from them. We review the state of the art in reproducing the Standard Model from string compactifications and summarize the lessons drawn from this fascinating quest. We describe novel scenarios and mechanisms that string theory provides to address some of the Standard Model puzzles as well as the most frequent signatures of new physics that could be detected in future experiments. We then comment on recent developments that connect, in a rather unexpected way, the Standard Model with quantum gravity and that may change our field theory notion of naturalness.
In physical networks trained using supervised learning, physical parameters are adjusted to produce desired responses to inputs. An example is an electrical contrastive local learning network of nodes connected by edges that adjust their conductances during training. When an edge conductance changes, it upsets the current balance of every node. In response, physics adjusts the node voltages to minimize the dissipated power. Learning in these systems is therefore a coupled double-optimization process, in which the network descends both a cost landscape in the high-dimensional space of edge conductances and a physical landscape—the power dissipation—in the high-dimensional space of node voltages. Because of this coupling, the physical landscape of a trained network contains information about the learned task. Here, we derive a structure-function relation for trained tunable networks and demonstrate that all the physical information relevant to the trained input-output relation can be captured by a tuning susceptibility, an experimentally measurable quantity. We supplement our theoretical results with simulations to show that the tuning susceptibility is correlated with functional importance and that we can extract physical insight into how the system performs the task from the conductances of highly susceptible edges. Our analysis is general and can be applied directly to mechanical networks, such as networks trained for protein-inspired function such as allostery.
Finding the transient and steady state properties of open quantum systems is a central problem in various fields of quantum technologies. Here, in this work, we present a quantum-assisted algorithm to determine the steady states of open system dynamics. By reformulating the problem of finding the fixed point of Lindblad dynamics as a feasibility semidefinite program, we bypass several well-known issues with variational quantum approaches to solving for steady states. We demonstrate that our hybrid approach allows us to estimate the steady states of higher dimensional open quantum systems and discuss how our method can find multiple steady states for systems with symmetries.
G4MP2 theory has proven to be a reliable and accurate quantum chemical composite method for the calculation of molecular energies using an approximation based on second-order perturbation theory to lower computational costs compared to G4 theory. However, it has been found to have significantly increased errors when applied to larger organic molecules with 10 or more nonhydrogen atoms. We report here on an investigation of the cause of the failure of G4MP2 theory for such larger molecules. One source of error is found to be the "higher-level correction (HLC)", which is meant to correct for deficiencies in correlation contributions to the calculated energies. This is because the HLC assumes that the contribution is independent of the element and the type of bonding involved, both of which become more important with larger molecules. We address this problem by adding an atom-specific correction, dependent on atom type but not bond type, to the higher-level correction. We find that a G4MP2 method that incorporates this modification of the higher-level correction, referred to as G4MP2A, becomes as accurate as G4 theory (for computing enthalpies of formation) for a test set of molecules with less than 10 nonhydrogen atoms as well as a set with 10-14 such atoms, the set of molecules considered here, with a much lower computational cost. The G4MP2A method is also found to significantly improve ionization potentials and electron affinities. Finally, we implemented the G4MP2A energies in a machine learning method to predict molecular energies.
We present a self-consistent algorithm for optimal control simulations of many-body quantum systems. The algorithm features a two-step synergism that combines discrete real-time machine learning (DRTL) with Quantum Optimal Control Theory (QOCT) using the time-dependent Schrödinger equation. Specifically, in step (1), DRTL is employed to identify a compact working space (i.e., the important portion of the Hilbert space) for the time evolution of the many-body quantum system in the presence of a control field (i.e., the initial or previously updated field), and in step (2), QOCT utilizes the DRTL-determined working space to find a newly updated control field for a chosen objective. Steps 1 and 2 are iterated until a self-consistent control objective value is reached such that the resulting optimal control field yields the same targeted objective value when the corresponding working space is systematically enlarged. Furthermore, to demonstrate this two-step self-consistent DRTL-QOCT synergistic algorithm, we perform optimal control simulations of strongly interacting 1D as well as 2D Heisenberg spin systems. In both scenarios, only a single spin (at the left end site for 1D and the upper left corner site for 2D) is driven by the time-dependent control fields to create an excitation at the opposite site as the target. It is found that, starting from all spin-down zero excitation states, the synergistic method is able to identify working spaces and convergence of the desired controlled dynamics with just a few iterations of the overall algorithm. In the cases studied, the dimensionality of the working space scales only quasi-linearly with the number of spins.
A combined large-scale first principles approach with machine learning and materials informatics is proposed to quickly sweep the chemistry-composition space of advanced high strength steels (AHSS). AHSS are composed of iron and key alloying elements such as aluminum and manganese. A systematic exploration of the distribution of aluminum and manganese atoms in iron is used to investigate low stacking fault energies configurations using first principles calculations. To overcome the computational cost of exploring the composition space, this process is sped up using an automated machine learning tool: DeepHyper. Here our results predict that it is energetically favorable for Al to stay away from a stacking fault, but Mn atoms do not affect the stacking fault energy and can stay in the vicinity of the fault. The distribution of Al and Mn atoms in systems containing stacking faults and the effects of their interactions on the equilibrium distribution are systematically analyzed.
Icosahedral boron materials, which include regular icosahedra of 12 boron atoms have gained increasing attention due to their potential applications as superhard materials, semiconductors, and energy storage media. However, the synthesis of high quality crystals of these materials has been a major barrier to the development of these applications. To enable computational prediction of synthesis conditions yielding high-quality icosahedral boron crystals, herein we tested and refined a set of ReaxFF parameters for the nucleation and growth of such crystals. We focused on matching the relative energies of small boron clusters obtained by density functional theory since such small clusters and similar motifs are likely present in crystal nuclei and at the interface of growing crystals. Using a training set of B 80 clusters, including a low-energy core–shell structure containing a B 12 icosahedron core and a high-energy single-shell structure produced in preliminary ReaxFF simulations, the ReaxFF parameter set was refined to better reproduce energies calculated by density functional theory (DFT). Among existing ReaxFF parameter sets and the machine-learning interatomic potentials MACE-MP-0, MACE-MP-0b3, MACE-MPA-0, PFP v7.0.0, and SevenNet-MF-ompa, only our new parameter set and PFP v7.0.0 correctly ranked these B 80 clusters. This refinement led to improved agreement with DFT for a test set of 58 clusters consisting of 8–103 boron atoms. Furthermore, our refined parameter set yielded greater local icosahedral structure than the previously existing ReaxFF parameter set for larger scale simulations of crystallization from supercooled liquid boron. Additionally, simulations of solid boron in contact with molten nickel using our refined ReaxFF parameters yielded a boron solubility value that agrees moderately well with experimental expectations, while the previous boron parameters gave a value that was much too low.
Machine learning methodologies can provide insight into Brønsted-Guggenheim-Scatchard specific ion interaction theory (SIT) parameter values where experimental data availability may be limited. This study develops and executes machine learning frameworks to model the SIT interaction coefficient, ε. Key findings include successful estimations of ε via artificial neural networks using clustering and value prediction approaches. Additionally, applicability to other chemical parameters is also assessed briefly. Models developed here provide support for a use-case of machine learning in geologic nuclear waste disposal research applications, namely in predictions of chemical behaviors of high ionic strength solutions (i.e., subsurface brines).
Density functional theory (DFT) stands as a cornerstone method in computational quantum chemistry and materials science due to its remarkable versatility and scalability. Yet, it suffers from limitations in accuracy, particularly when dealing with strongly correlated systems. To address these shortcomings, recent work has begun to explore how machine learning can expand the capabilities of DFT: an endeavor with many open questions and technical challenges. In this work, we present GradDFT a fully differentiable JAX-based DFT library, enabling quick prototyping and experimentation with machine learning-enhanced exchange–correlation energy functionals. GradDFT employs a pioneering parametrization of exchange–correlation functionals constructed using a weighted sum of energy densities, where the weights are determined using neural networks. Moreover, GradDFT encompasses a comprehensive suite of auxiliary functions, notably featuring a just-in-time compilable and fully differentiable self-consistent iterative procedure. To support training and benchmarking efforts, we additionally compile a curated dataset of experimental dissociation energies of dimers, half of which contain transition metal atoms characterized by strong electronic correlations. The software library is tested against experimental results to study the generalization capabilities of a neural functional across potential energy surfaces and atomic species, as well as the effect of training data noise on the resulting model accuracy.
We apply machine-learning techniques to the effective-field-theory analysis of the e + e − → W + W − processes at future lepton colliders, and demonstrate their advantages in comparison with conventional methods, such as optimal observables. In particular, we show that machine-learning methods are more robust to detector effects and backgrounds, and could in principle produce unbiased results with sufficient Monte Carlo simulation samples that accurately describe experiments. This is crucial for the analyses at future lepton colliders given the outstanding precision of the e + e − → W + W − measurement (~ 10−4 in terms of anomalous triple gauge couplings or even better) that can be reached. Our framework can be generalized to other effective-field-theory analyses, such as the one of e + e − → t t ¯ or similar processes at muon colliders.
Predicting the structural properties of water and simple fluids confined in nanometer scale pores and channels is essential in, for example, energy storage and biomolecular systems. Classical continuum theories fail to accurately capture the interfacial structure of fluids. In this work, we develop a deep learning-based quasi-continuum theory (DL-QT) to predict the concentration and potential profiles of a Lennard-Jones (LJ) fluid and water confined in a nanochannel. The deep learning model is built based on a convolutional encoder–decoder network (CED) and is applied for high-dimensional surrogate modeling to relate the fluid properties to the fluid–fluid potential. The CED model is then combined with the interatomic potential-based continuum theory to determine the concentration profiles of a confined LJ fluid and confined water. Further, we show that the DL-QT model exhibits robust predictive performance for a confined LJ fluid under various thermodynamic states and for water confined in a nanochannel of different widths. The DL-QT model seamlessly connects molecular physics at the nanoscale with continuum theory by using a deep learning model.
The high brightness, low emittance electron beams achieved in modern X-ray free-electron lasers (XFELs) have enabled powerful X-ray imaging tools, allowing molecular systems to be imaged at picosecond time scales and sub-nanometer length scales. One of the most promising directions for increasing the brightness of XFELs is through the development of novel photocathode materials. Whereas past efforts aimed at discovering photocathode materials have typically employed trial-and-error-based iterative approaches, this work represents the first data-driven screening for high brightness photocathode materials. Through screening over 74 000 semiconducting materials, a vast photocathode dataset is generated, resulting in statistically meaningful insights into the nature of high brightness photocathode materials. This screening results in a diverse list of photocathode materials that exhibit intrinsic emittances that are up to 4x lower than currently used photocathodes. In a second effort, multiobjective screening is employed to identify the family of M 2 O (M = Na, K, Rb) that exhibits photoemission properties that are comparable to the current state-of-the-art photocathode materials, but with superior air stability. This family represents perhaps the first intrinsically bright, visible light photocathode materials that are resistant to reactions with oxygen, allowing for their transport and storage in dry air environments.
Not provided.