Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “activation function”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Multiobjective Hyperparameter Optimization for Deep Learning Interatomic Potential Training Using NSGA-II

Deep neural network (DNN) potentials are an emerging tool for simulation of dynamical atomistic systems, with the promise of quantum mechanical accuracy at speedups of 10000$\times$. As with other DNN methods, hyperparameters used during training can make a substantial difference in model accuracy, and optimal settings vary with dataset. To enable rapid tuning of hyperparameters for DNN potential training, we developed a scalable multiobjective optimization evolutionary algorithm for supercomputers and tested it on the Summit system at the Oak Ridge Leadership Computing Facility (OLCF). The multiobjective approach is required due to the coupling of two learned values defining the potential: the energy and force. Using a large-scale implementation of the NSGA-II algorithm adapted for training DNN potentials, we discovered several optimal multiobjective combinations, including best choices of activation functions, learning rate scaling scheme, and pairing of the two radial cutoffs used in the three dimensional descriptor function.

Coletti, Mark↗

Emulator-Based Bayesian Calibration of the CISNET Colorectal Cancer Models

Purpose To calibrate Cancer Intervention and Surveillance Modeling Network (CISNET)'s SimCRC, MISCAN-Colon, and CRC-SPIN simulation models of the natural history colorectal cancer (CRC) with an emulator-based Bayesian algorithm and internally validate the model-predicted outcomes to calibration targets.Methods We used Latin hypercube sampling to sample up to 50,000 parameter sets for each CISNET-CRC model and generated the corresponding outputs. We trained multilayer perceptron artificial neural networks (ANNs) as emulators using the input and output samples for each CISNET-CRC model. We selected ANN structures with corresponding hyperparameters (i.e., number of hidden layers, nodes, activation functions, epochs, and optimizer) that minimize the predicted mean square error on the validation sample. We implemented the ANN emulators in a probabilistic programming language and calibrated the input parameters with Hamiltonian Monte Carlo-based algorithms to obtain the joint posterior distributions of the CISNET-CRC models' parameters. We internally validated each calibrated emulator by comparing the model-predicted posterior outputs against the calibration targets.Results The optimal ANN for SimCRC had 4 hidden layers and 360 hidden nodes, MISCAN-Colon had 4 hidden layers and 114 hidden nodes, and CRC-SPIN had 1 hidden layer and 140 hidden nodes. The total time for training and calibrating the emulators was 7.3, 4.0, and 0.66 h for SimCRC, MISCAN-Colon, and CRC-SPIN, respectively. The mean of the model-predicted outputs fell within the 95% confidence intervals of the calibration targets in 98 of 110 for SimCRC, 65 of 93 for MISCAN, and 31 of 41 targets for CRC-SPIN.Conclusions Using ANN emulators is a practical solution to reduce the computational burden and complexity for Bayesian calibration of individual-level simulation models used for policy analysis, such as the CISNET CRC models. In this work, we present a step-by-step guide to constructing emulators for calibrating 3 realistic CRC individual-level models using a Bayesian approach.

artificial neural networks↗

Biophysical characterization and a roadmap towards the NMR solution structure of G0S2, a key enzyme in non-alcoholic fatty liver disease

In the United States non-alcoholic fatty liver disease (NAFLD) is the most common form of chronic liver disease, affecting an estimated 80 to 100 million people. It occurs in every age group, but predominantly in people with risk factors such as obesity and type 2 diabetes. NAFLD is marked by fat accumulation in the liver leading to liver inflammation, which may lead to scarring and irreversible damage progressing to cirrhosis and liver failure. In animal models, genetic ablation of the protein G0S2 leads to alleviation of liver damage and insulin resistance in high fat diets. The research presented in this paper aims to aid in rational based drug design for the treatment of NAFLD by providing a pathway for a solution state NMR structure of G0S2. Here we describe the expression of G0S2 in an E. coli system from two different constructs, both of which are confirmed to be functionally active based on the ability to inhibit the activity of Adipose Triglyceride Lipase. In one of the constructs, preliminary NMR spectroscopy measurements show dominant alpha-helical characteristics as well as resonance assignments on the N-terminus of G0S2, allowing for further NMR work with this protein. Additionally, the characterization of G0S2 oligomers are outlined for both constructs, suggesting that G0S2 may defensively exist in a multimeric state to protect and potentially stabilize the small 104 amino acid protein within the cell. This information presented on the structure of G0S2 will further guide future development in the therapy for NAFLD.

59 BASIC BIOLOGICAL SCIENCES↗

Understanding the Inner-Workings of Language Models Through Representation Dissimilarity

We use model stitching to understand the internal representations of language models. Similar to vision models, we find that "more is better," and representations learned with more data and larger width can improve the performance of weaker models via stitching. We likewise find that certain architecture choices, using GeLU vs SoLU activation functions, influence the quality of learned representations. Finally, model stitching (as opposed to other model diagnostic methods, like mode connectivity) can localize the different generalization strategies of text classifiers under domain shift to certain hidden layers.

Brown, Davis R.↗

Toward a Photomagnetic Mechanism for f-element separations (Final Technical Report)

The intent behind this proposal was to evaluate magnetic field effects of rare-earth ions on radical pairs. Rare-earth ions are notoriously challenging to separate based on similar thermodynamic characteristics (e.g. solubility) or ion radius, whereas magnetism varies starkly from rare-earth to rare-earth, even when ions are immediately adjacent on the periodic table. Magnetism could be a useful property then for separations, but a mechanism that is effective and selective for magnetic characteristics must be produced. A reactive mechanism for separations could likewise be interesting for selective separations, if a magnetic handle for such a separation could be defined. This proposal sought to understand a specific reactive, magnetic handle for a reactive separations scheme by creating proof of concept model complexes that place rare earths in proximity to photochemically active functional groups.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BiomeProbe [Abstract]

We propose the development and application of a microtiter plate-based chemical probe assay that can facilitate the rapid quantification of fecal microbiome-associated enzyme functional activities (“gut pharmacomicrobiomics”). Our initial proof-of-concept project with FedImpact is to show that one or more glucuronidase probes can be used for sensitive and reproducible characterization of fecal microbe glucuronidase activities. Glucuronidase enzymes are important for contributing to enterohepatic recycling of therapeutic drug compounds, resulting in impaired or prolonged drug activity. Importantly, the assay to be developed will permit facile ‘swapping’ of chemical probes to characterize different and multiple enzyme activities. Similarly, the assay is naïve to the biological sample to be analyzed.

59 BASIC BIOLOGICAL SCIENCES↗

Domain Aware Deep-learning Algorithms Integrated with Scientific-computing Technologies (DADAIST)

This technical report summarized the contribution of the DADAIST project funded by the Data Model Convergence Initiative via the Laboratory Directed Research and Development (LDRD) investments at Pacific Northwest National Laboratory (PNNL). Specifically, we report the development of the NeuroMANCER (Neural Modules with Adaptive Nonlinear Constraints and Efficient Regularizations), a new open-source Scientific Machine Learning library for formulating and solving parametric constrained optimization problems, physics-informed system identification, and parametric optimal control problems. NeuroMANCER is using differentiable programming to combine modern data-driven models and optimization modeling language into a coherent algorithmic and software framework. NeuroMANCER is a Pytorch-based framework and adopts much of its philosophy focused on research and development, rapid prototyping, and streamlined deployment. Strong emphasis is given to extensibility, interoperability with the PyTorch ecosystem, and quick adaptability to custom domain problems. Neuromancer repository contains a comprehensive library of differentiable modules, including custom activation functions, matrix factorizations, deep learning architectures, neural differential equations, differential equation solvers, implicit layers such as iterative solvers, high-level API for symbolic expressions, API for modeling and control of dynamical systems, and extensive set of tutorial code examples in the form of python scripts and jupyter notebooks.

97 MATHEMATICS AND COMPUTING↗

Leveraging Artificial Intelligence to Predict Novel Eutectic Alloys

The goal of this project was to train an artificial neural network (ANN) to predict the fractional composition and melting point of eutectic alloys using fundamental atomic properties as inputs. The fundamental properties considered include atomic number, atomic weight, atomic radius, valence electron concentration, electronegativity, and electron affinity. The project involved several phases, starting with data preparation, where phase diagram data was harvested from the ASM International database. Approximately 1300 binary eutectics were collected and cleaned to ensure relevance and accuracy. A regression model was selected for training, utilizing a rectified linear unit as the activation function. Various model configurations were evaluated for predictive accuracy, with validation techniques employed to ensure robustness. The model demonstrated predictive capabilities above random guessing and was able to achieve up to 11% accuracy under certain conditions. An ablative test identified atomic radius and valence electron concentration as critical inputs for model performance. Incorporating the melting point of atomic constituents improved accuracy significantly, although ultimately the model’s predictive capability still fell short of the 80% target. This report details the methodology, results, and implications of the research, contributing to the understanding of employing artificial intelligence to predict the phase transition behavior of eutectic alloys.

36 MATERIALS SCIENCE↗

Early Drug Discovery and Development of Novel Cancer Therapeutics Targeting DNA Polymerase Eta (POLH)

Polymerase eta (or Pol η or POLH) is a specialized DNA polymerase that is able to bypass certain blocking lesions, such as those generated by ultraviolet radiation (UVR) or cisplatin, and is deployed to replication foci for translesion synthesis as part of the DNA damage response (DDR). Inherited defects in the gene encoding POLH (a.k.a., XPV) are associated with the rare, sun-sensitive, cancer-prone disorder, xeroderma pigmentosum, owing to the enzyme’s ability to accurately bypass UVR-induced thymine dimers. In standard-of-care cancer therapies involving platinum-based clinical agents, e.g., cisplatin or oxaliplatin, POLH can bypass platinum-DNA adducts, negating benefits of the treatment and enabling drug resistance. POLH inhibition can sensitize cells to platinum-based chemotherapies, and the polymerase has also been implicated in resistance to nucleoside analogs, such as gemcitabine. POLH overexpression has been linked to the development of chemoresistance in several cancers, including lung, ovarian, and bladder. Co-inhibition of POLH and the ATR serine/threonine kinase, another DDR protein, causes synthetic lethality in a range of cancers, reinforcing that POLH is an emerging target for the development of novel oncology therapeutics. Using a fragment-based drug discovery approach in combination with an optimized crystallization screen, we have solved the first X-ray crystal structures of small novel drug-like compounds, i.e., fragments, bound to POLH, as starting points for the design of POLH inhibitors. The intrinsic molecular resolution afforded by the method can be quickly exploited in fragment growth and elaboration as well as analog scoping and scaffold hopping using medicinal and computational chemistry to advance hits to lead. An initial small round of medicinal chemistry has resulted in inhibitors with a range of functional activity in an in vitro biochemical assay, leading to the rapid identification of an inhibitor to advance to subsequent rounds of chemistry to generate a lead compound. Importantly, our chemical matter is different from the traditional nucleoside analog-based approaches for targeting DNA polymerases.

60 APPLIED LIFE SCIENCES↗

An Accuracy-Maximization Approach for Claims Classifiers in Document Content Analytics for Cybersecurity

This paper presents our research approach and findings towards maximizing the accuracy of our classifier of feature claims for cybersecurity literature analytics, and introduces the resulting model ClaimsBERT. Its architecture, after extensive evaluations of different approaches, introduces a feature map concatenated with a Bidirectional Encoder Representation from Transformers (BERT) model. We discuss deployment of this new concept and the research insights that resulted in the selection of Convolution Neural Networks for its feature mapping aspects. We also present our results showing ClaimsBERT to outperform all other evaluated approaches. This new claims classifier represents an essential processing stage within our vetting framework aiming to improve the cybersecurity of industrial control systems (ICS). Furthermore, in order to maximize the accuracy of our new ClaimsBERT classifier, we propose an approach for optimal architecture selection and determination of optimized hyperparameters, in particular the best learning rate, number of convolutions, filter sizes, activation function, the number of dense layers, as well as the number of neurons and the drop-out rate for each layer. Fine-tuning these hyperparameters within our model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original model to a 97% accuracy obtained with ClaimsBERT.

Ameri, Kimia (ORCID:0000000328791871)↗

In Vitro Selection of Antibodies Targeting Yersinia pestis Membrane Lipids Using Nanodisc-Based Antigen Presentation

Proteins are the most common targets for antibody discovery and vaccine development, but their sequence variability can limit the breadth of resulting antigens. Lipids represent an alternative class of antigens due to their structural conservation and roles in host–pathogen interactions. Here, we describe the development and optimization of an in vitro antibody selection workflow using lipid-containing nanodiscs as antigen presentation platforms to enable phage and yeast display selections under conditions adapted for these non-protein targets. Lipopolysaccharide (LPS) nanodiscs were first used as a model system to evaluate selection strategies, including competitive and subtractive approaches to reduce non-specific binders, yielding peptide and single-chain variable fragment (scFv) binders that were affinity matured to improve binding signals. The same approach was subsequently used to select scFv antibodies that recognize lipid nanodiscs prepared from Yersinia pestis membrane lipid extracts. These antibodies show binding to lipid nanodiscs derived from Y. pestis, with evidence of selectivity relative to control nanodiscs. Overall, this work establishes a workflow for antibody selection against lipid-containing nanodisc antigens and highlights practical considerations associated with these targets. The approach may be useful for generating affinity reagents to membrane-associated lipids, although further characterization is required to define antigen specificity and functional activity.

59 BASIC BIOLOGICAL SCIENCES↗

GAINN: The Galaxy Assembly and Interaction Neural Networks for High-redshift JWST Observations

We present the Galaxy Assembly and Interaction Neural Networks (Gainn), a series of artificial neural networks for predicting the redshift, stellar mass, halo mass, and mass-weighted age of simulated galaxies based on James Webb Space Telescope (JWST) photometry. Our goal is to determine the best neural network for predicting these variables at 11 < z < 15. The parameters of the optimal neural network can then be used to estimate these variables for real, observed galaxies. The inputs of the neural networks are JWST filter magnitudes of a subset of five broadband filters (F150W, F200W, F277W, F356W, and F444W) and two medium-band filters (F162M and F182M). We compare the performance of the neural networks using different combinations of these filters, as well as different activation functions and numbers of layers. The best neural network predicted redshift with a normalized rms error of $0.010^{+0.003}_{-0.001}$, stellar mass with rms = $0.089^{+0.044}_{-0.022}$, halo mass with a mean-squared error of $0.022^{+0.014}_{-0.008}$, and mass-weighted age with rms = $12.466^{+5.065}_{-2.408}$. We also test the performance of Gainn on real data from MACS0647JD, an object observed by JWST. Predictions from Gainn for the first projection of the object (JD1) have normalized bias $\langle$Δz$\rangle$ < 0.00228, which is significantly smaller than found with template-fitting methods. We find that the optimal filter combination is F277W, F356W, F162M, and F200W when considering both theoretical accuracy and observational resources from JWST.

97 MATHEMATICS AND COMPUTING↗

A Deep Learning Modeling Framework to Capture Mixing Patterns in Reactive-Transport Systems

Prediction and control of chemical mixing are vital for many scientific areas such as subsurface reactive transport, climate modeling, combustion, epidemiology, and pharmacology. Due to the complex nature of mixing in heterogeneous and anisotropic media, the mathematical models related to this phenomenon are not analytically tractable. Numerical simulations often provide a viable route to predict chemical mixing accurately. However, contemporary modeling approaches for mixing cannot utilize available spatial-temporal data to improve the accuracy of the future prediction and can be compute-intensive, especially when the spatial domain is large and for long-term temporal predictions. To address this knowledge gap, in this work we will present in this paper a deep learning (DL) modeling framework applied to predict the progress of chemical mixing under fast bimolecular reactions. This framework uses convolutional neural networks (CNN) for capturing spatial patterns and long short-term memory (LSTM) networks for forecasting temporal variations in mixing. By careful design of the framework—placement of non-negative constraint on the weights of the CNN and the selection of activation function, the framework ensures non-negativity of the chemical species at all spatial points and for all times. Our DL-based framework is fast, accurate, and requires minimal data for training. The time needed to obtain a forecast using the model is a fraction (≈ O(-6)) of the time needed to obtain the result using a high-fidelity simulation. To achieve an error of 10% (measured using the infinity norm) for capturing local-scale mixing features such as interfacial mixing, only 24% to 32% of the sequence data for model training is required. To achieve the same level of accuracy for capturing global-scale mixing features, the sequence data required for model training is 64% to 70% of the total spatial-temporal data. Hence, the proposed approach—a fast and accurate way to forecast long-time spatial-temporal mixing patterns in heterogeneous and anisotropic media—will be a valuable tool for modeling reactive-transport in a wide range of applications.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Development of a Deep Learning Model for Predicting the Drag Coefficients of Spherical and Non-Spherical Particles,

There is yet to be a well-established drag model for non-spherical particles required in a particle-laden flow that could cover a wide range of sphericities. This talk will explore the development of a general drag model for non-spherical particles by applying deep learning using available experimental data available in the literature. The integration of several raw experimental measurements from different sources and research directions allows the training of robust Artificial Intelligence and Machine Learning (ML) models. Neural networks are an ML approach inspired by the inner biological workings of the brain. This work aims to develop a Deep Neural Network (DNN) that predicts drag coefficient values with the ability to adapt appropriately to unseen data. Given the limited number of data points available and the variance found within the data collected from various sources, challenges may arise when looking to train the model. Our study tests and implements various model regularization techniques and assesses different loss and activation functions for the proposed DNN. The proposed model considers a broader range of features other than sphericity and Reynold number. These features include density ratio, solid volume fraction, lengthwise and crosswise sphericity, and more. Furthermore, we present the features that play a significant role in predicting different drag coefficients through feature importance. Within the investigated parameter ranges in this study, the following conclusions can be achieved and summarized below: • An improved drag coefficient model can be developed by considering more features such as, aspect ratio, lengthwise sphericity, crosswise sphericity, and density ratio. • DNN model can predict better results compared to traditional methods using MAE metric. • The proposed model addresses data challenges such as limited data and extreme data points through expanded feature-set and regularization. • Three major features that mostly affect the drag coefficient were identified from a feature importance analysis.

Presa-Reyes, Maria↗

Project Final Report - Advanced Oxygen Separation from Air Using a Novel Mixed Matrix Membrane

This report describes a summary of the findings of the project. The technical work was divided into three activities: 1) technoeconomic analysis; 2) membrane component development (support layer, selective layer development, and gutter layer formulation); and 3) hollow fiber formation. The technoeconomic analysis includes operating, capital, and materials costs of an oxygen enrichment system with 1-3 stages. The analysis concluded that the desired gas quality (90-95% O2) cannot be reached with a single stage. The triple stage system yielded the best overall product quality at lower membrane performance levels but with the highest cost. For the membrane development activities, functionalized nanodiamonds have been shown to slightly improve membrane formation behaviors and gas permeability performance. For example, support layers based on polysulfone-nanodiamond composites have been formed that increase the material’s O2 permeability by over a factor of 150, however, at the cost of selectivity that suggests a change in transport mechanism. Furthermore, nanodiamonds have been found to improve membrane durability. Improvements in selective and gutter layers also have been accomplished. An additional accomplishment is the demonstration of phase inversion of polysulfone-nanodiamond composites that ensures that hollow fiber formation can be performed. Finally, the interfaces between support, gutter, and selective layers were briefly examined and the resulting membrane material consisting of all three layers gave good performance in an O2 separation from N2. The three-layer prototype membrane gave permeability of 29.2-32.2 Barrers, which equates to permeance of approximately 290-320 GPU, which is lower than the TEA informed goal of 500 GPU. O2/N2 selectivity ranged from 5.2-6.2, which is consistent with the less costly two-stage implementation of the technology. In summary, new materials have been developed for O2 concentration from air and are well-suited to a hollow fiber format.

20 FOSSIL-FUELED POWER PLANTS↗

Efficient Implementation of Artificial Neural Networks for Sensor Data Analysis Based on a Genetic Algorithm

The reliability of many industrial processes depends on the sensor system. However, these sensors can be affected by noise, perturbations and failures. Hence, sensor monitoring and diagnosis are fundamental to guarantee the quality of an industrial process. Nowadays, artificial neural networks (ANN) are widely used in sensor signal processing and diagnosis. However, those ANNs usually require many artificial neurons, being difficult to implement in software and hardware due to their high computational costs. This paper presents an optimized implementation of artificial neurons in ANNs for sensor data analysis using a Genetic Algorithm (GA). The objective of GA is to find an adequate segmentation to reduce the activation function approximation error. One of the advantages of the proposed approach is that the cost function used in GA considers the effect of factors such as the ANN architecture or the number of bits used in arithmetic operations. The proposed ANN implementation technique aims to get the best possible approximation for a specific ANN architecture, making easier its implementation in software and hardware. Simulation and experimental results using FPGA (Field Programmable Gate Array) prove the advantages of the proposed approach for implementing sensor data analysis systems based on ANNs.

D estefani, André↗

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. DMMs using deep neural networks to parametrize the transition of Markov probability distributions have recently been shown to provide more expressiveness in modeling sequential data and dynamical system responses. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a rigorous analytical method to prove the necessary and sufficient conditions of DMM's stochastic stability. This task is achieved by spectral analysis of the efficiently computed Jacobians of probabilistic maps modeled by deep neural networks. We make theoretical connections between the eigenvalues of neural network's weights and the different activation function types used on the stability and overall dynamic behavior of DMMs with Gaussian distributions. We empirically substantiate our theoretical results on stochastic stability and eigenvalue spectra via several numerical experiments. Formal stability guarantees of DMMs can substantially improve their robustness and trustworthiness, necessary for reliable use in safety-critical real-world applications.

Drgona, Jan↗

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a novel stability analysis method and provide sufficient conditions of DMM's stochastic stability. The proposed stability analysis is based on the contraction of probabilistic maps modeled by deep neural networks. We make connections between the spectral properties of neural network's weights and different types of used activation function on the stability and overall dynamic behavior of DMMs with Gaussian distributions. Based on the theory, we propose a few practical methods for designing constrained DMMs with guaranteed stability. We empirically substantiate our theoretical results via intuitive numerical experiments using the proposed stability constraints.

Drgona, Jan↗