Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “random number generators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Constructing high complexity synthetic libraries of long ORFs using in vitro selection

We present a method that can significantly increase the complexity of protein libraries used for in vitro or in vivo protein selection experiments. Protein libraries are often encoded by chemically synthesized DNA, in which part of the open reading frame is randomized. There are, however, major obstacles associated with the chemical synthesis of long open reading frames, especially those containing random segments. Insertions and deletions that occur during chemical synthesis cause frameshifts, and stop codons in the random region will cause premature termination. These problems can together greatly reduce the number of full-length synthetic genes in the library. We describe a strategy in which smaller segments of the synthetic open reading frame are selected in vitro using mRNA display for the absence of frameshifts and stop codons. These smaller segments are then ligated together to form combinatorial libraries of long uninterrupted open reading frames. This process can increase the number of full-length open reading frames in libraries by up to two orders of magnitude, resulting in protein libraries with complexities of greater than 10(13). We have used this methodology to generate three types of displayed protein library: a completely random sequence library, a library of concatemerized oligopeptide cassettes with a propensity for forming amphipathic alpha-helical or beta-strand structures, and a library based on one of the most common enzymatic scaffolds, the alpha/beta (TIM) barrel. Copyright 2000 Academic Press.

Non-NASA Center↗

Probability techniques for reliability analysis of composite materials

Traditional design approaches for composite materials have employed deterministic criteria for failure analysis. New approaches are required to predict the reliability of composite structures since strengths and stresses may be random variables. This report will examine and compare methods used to evaluate the reliability of composite laminae. The two types of methods that will be evaluated are fast probability integration (FPI) methods and Monte Carlo methods. In these methods, reliability is formulated as the probability that an explicit function of random variables is less than a given constant. Using failure criteria developed for composite materials, a function of design variables can be generated which defines a 'failure surface' in probability space. A number of methods are available to evaluate the integration over the probability space bounded by this surface; this integration delivers the required reliability. The methods which will be evaluated are: the first order, second moment FPI methods; second order, second moment FPI methods; the simple Monte Carlo; and an advanced Monte Carlo technique which utilizes importance sampling. The methods are compared for accuracy, efficiency, and for the conservativism of the reliability estimation. The methodology involved in determining the sensitivity of the reliability estimate to the design variables (strength distributions) and importance factors is also presented.

Wetherhold, Robert C.↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Neural MUSE Analysis

Researchers at Oak Ridge National Laboratory (ORNL) created data as part of the MUSE (Multi-Agency Urban Search Experiment Detector and Algorithm Test Bed) project simulating illicit nuclear materials located in various buildings along a road. In the simulation, a truck containing a radiation detector drives down the road gathering listmode data (counting the and energy of incident gamma radiation). Building materials, source shielding, driving speed, truck direction, truck location on the road, source type, and source placement are all varied between runs of the data set. This data was created using deterministic neutron transport and Monte Carlo methods through a combination of SCALE, MAVRIC, MCNP, and GADRAS. As part of a follow-on NA-22 project, two Kaggle competitions were created to determine the best algorithms for finding and identifying gamma sources in this simulated urban environment. The winning algorithm was neural network-based and had a test accuracy of 76.4% accuracy for source identification. This work seeks to build upon this work and improve the results through the application of novel machine learning techniques. As a first step, the data was classified by a simple Convolutional Neural Network (CNN) To accomplish this, the data was first preprocessed into “waterfall plots.” These plots are composed of energy vs count plots that are stacked vertically to show progression in time. The horizontal axis indicating the particle energy incorporated user defined bin spacing with options for in linear-, logarithmic-, square root-, and user-spaced bins. The z or color dimension showed the number of counts corresponding the energy-time combination. This data was then used to generate more data, by generating a local estimate of the mean of the distribution for a bin and then randomly re-sampling that bin from a Poisson distribution. Once all of this data was generated, it was fed into a well-known CNN architecture, ResNet50. The output layer of this model was removed and replaced with layers corresponding to the shape desired isotope outputs. The provided training data was used to train the classifier and the remaining testing data was used to evaluate the model. Results are soon to be forthcoming.

61 RADIATION PROTECTION AND DOSIMETRY↗

Efficient generation of statistically good pseudonoise by linearly interconnected shift registers

A number of efficient new algorithms for generating digital pseudonoise are presented. The algorithms are efficient because a large number of new pseudorandom bits are generated at each iteration of a computer implementation or at each clock pulse of a hardware implementation. The maximal-length shift register sequences obtained have good randomness properties. Some of the properties of p-n sequences are reviewed and the results of an extensive statistical evaluation of the pseudonoise are presented.

Hurd, W. J.↗

FIRM image analysis: A machine learning workflow for quantifying extracellular matrix components from electron microscopy images

The extracellular matrix (ECM) is a complex network of biomolecules that plays an integral role in the structure, processes, and signaling mechanisms of cells and tissues. Identifying and quantifying changes in these matrix components provides insight into the mechanisms behind specific tissue remodeling processes; however, quantifying these changes is challenging due to difficult imaging conditions, complexity of the ECM, and the subtlety of these changes. Current imaging techniques allow us to visualize these critical remodeling events and developments in image analysis have employed a combination of analysis software and machine learning techniques to improve the efficiency and accuracy with which features are measured. Although image analysis has seen much improvement in recent years, there has been no technique developed to address ambiguity in feature edges in electron microscopy images. Presented here is a new machine learning-based workflow for the analysis of microscopy images named FIRM (Feature Identification from Raw Microscopy) that uses a random forest classifier to identify ECM features of interest and generate binary segmentation masks for quantification with ImageJ-FIJI. FIRM performed with an F1 score of 0.794 and greater than 80% accuracy for number and size of features detected. FIRM had similar deviation from the ground truth in the number of identified fibrils, fibril size, and size distributions when compared to human analyses. The results suggest that FIRM performs as well as manual analysis and requires a fraction of the time. This analysis technique is more efficient, eliminates user bias, and can be easily optimized to identify a variety of features, making it useful for any discipline requiring image analysis.

Science & Technology - Other Topics↗

COWALKER:EFFECTIVE TRANSPORT PROPERTIES OF COMPOSITE MATERIALS

SF-23-026 This software computes effective transport properties of composite materials involving fibers and nanoparticles using a random-walk algorithm that efficiently scales to an arbitrary number of processes and cores. Effective transport properties (thermal, electrical) are key to bridge the microstructure of complex materials with its macroscopic behavior. Traditional approaches either use effective medium approximations (closed mathematical expressions that are approximation for certain conditions) or continuum simulation models such as finite element or finite volume, which require the generation of a mesh for each configuration explored. cowalker leverages the equivalence between laplacian or heat equation-based models and random walks to compute the asymptotic transport properties from an ensemble of first sojourn times of a random walker moving through the composite material. This allows us to directly define a composite material as a collection of particles and use algorithms developed for molecular dynamics to quickly compute the intersection of the walker with the different interfaces in the material. cowalker is developed in C++, and it relies on the GNU Scientific Library for random generation. cowalker is currently delivered as source code, so the GSL library is not included in cowalker's distribution. A more userfriendly version, cowalker.jl is currently in development and will be released as part of cowalker.

YANGUAS-GIL, ANGEL↗

The simulated depth history of dust grains in the lunar regolith

Trajectories giving the individual depth variations of lunar dust grains with time are randomly generated by a Monte Carlo code, where the variables are the mass and speed distribution of meteorites at the lunar surface and the geometrical shape of impact craters. A statistical analysis of a great number of such trajectories is then used to define the 'average' depth history of lunar dust grains for two grain radii: 1 and 50 microns. This yields: (1) estimates for time constants involved during the dynamic evolution of the regolith, (2) a model for the layering of the regolith, and (3) a better understanding of the basic dust-grain mechanisms responsible for the formation of the most mature lunar soil samples. The validity of various soil models proposed for the dynamic evolution of the regolith is discussed in terms of experimental constraints on the models.

Duraud, J. P.↗

Stochastic modal velocity field in rough-wall turbulence

Stochastically generated instantaneous velocity profiles are used to reproduce the outer region of rough-wall turbulent boundary layers in a range of Reynolds numbers extending from the wind tunnel to field conditions. Each profile consists in a sequence of steps, defined by the modal velocities and representing uniform momentum zones (UMZs), separated by velocity jumps representing the internal shear layers. Height-dependent UMZ is described by a minimal set of attributes: thickness, mid-height elevation, and streamwise (modal) and vertical velocities. These are informed by experimental observations and reproducing the statistical behaviour of rough-wall turbulence and attached eddy scaling, consistent with the corresponding experimental datasets. Sets of independently generated profiles are reorganized in the streamwise direction to form a spatially consistent modal velocity field, starting from any randomly selected profile. The operation allows one to stretch or compress the velocity field in space, increases the size of the domain and adjusts the size of the largest emerging structures to the Reynolds number of the simulated flow. By imposing the autocorrelation function of the modal velocity field to be anchored on the experimental measurements, we obtain a physically based spatial resolution, which is employed in the computation of the velocity spectrum, and second-order structure functions. The results reproduce the Kolmogorov inertial range extending from the UMZ and their attached-eddy vertical organization to the very-large-scale motions (VLSMs) introduced with the reordering process. The dynamic role of VLSM is confirmed in the –u'w' co-spectra and in their vertical derivative, representing a scale-dependent pressure gradient contribution.

42 ENGINEERING↗

A Synthetic Transcription Factor and Core Promoter System in Picochlorum renovo Enables Tunable Gene Expression

Picochlorum renovo is a recently characterized microalga of industrial interest. Its rapid growth rate, and high temperature and salinity tolerances make P. renovo an attractive candidate for industrial scale cultivation and downstream production of sustainable fuels and chemicals. Currently, genetic tools for many non-model microalgae are limited and would greatly benefit from an orthogonal gene expression system to bypass host regulation. Additionally, the engineering of complex metabolic pathways in eukaryotic organisms to optimize growth or biosynthesize high value products often requires tunable expression of each gene in a pathway. Here we explore a tunable orthogonal gene expression system using a synthetic transcription factor (sTF) and core promoters (CPs) conferring expression of the fluorescent protein mCherry to quantify protein expression. The sTF paired with the relevant binding site (BS) led to an ~5X increase in reporter gene expression compared to the native RuBisCo promoter, however had limited tunability with increasing BS number. Quantification of mCherry expression under 34 different CPs paired with the sTF and BS showed an order of magnitude of expression tunability. Future work with this system will entail generation of an overexpression library via random integration of the relevant BS in an sTF expressing P. renovo strain. With this sTF and CP system we aim to greatly improve growth rates and product titers in photosynthetic organisms, while also providing a potentially universal gene expression system for microalgae.

algae↗

Particle-number distribution in large fluctuations at the tip of branching random walks

Here, we investigate properties of the particle distribution near the tip of one-dimensional branching random walks at large times t , focusing on unusual realizations in which the rightmost lead particle is very far ahead of its expected position, but still within a distance smaller than the diffusion radius ~$\sqrt{t}$. Our approach consists in a study of the generating function $G_{Δx}(λ) = Σ_n$ ${λ^n}p_n(Δx)$ for the probabilities $p_n(Δx)$ of observing $\textit{n}$ particles in an interval of given size $Δ\textit{x}$ from the lead particle to its left, fixing the position of the latter. This generating function can be expressed with the help of functions solving the Fisher-Kolmogorov-Petrovsky-Piscounov (FKPP) equation with suitable initial conditions. In the infinite-time and large-$Δ\textit{x}$ limits, we find that the mean number of particles in the interval grows exponentially with $Δ\textit{x}$, and that the generating function obeys a nontrivial scaling law, depending on $Δ\textit{x}$ and λ through the combined variable $[Δx — f(λ)]^3 / Δx^2$, where $\textit{f}$(λ) ≡ – ln(1 – λ) – ln [– ln(1 – λ)]. From this property, one may conjecture that the growth of the typical particle number with the size of the interval is slower than exponential, but, surprisingly enough, only by a subleading factor at large Δ$\textit{x}$. The scaling we argue is consistent with results from a numerical integration of the FKPP equation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The reproduction number and its probability distribution for stochastic viral dynamics

We consider stochastic models of individual infected cells. The reproduction number, R, is understood as a random variable representing the number of new cells infected by one initial infected cell in an otherwise susceptible (target cell) population. Variability in R results partly from heterogeneity in the viral burst size (the number of viral progeny generated from an infected cell during its lifetime), which depends on the distribution of cellular lifetimes and on the mechanism of virion release. We analyse viral dynamics models with an eclipse phase: the period of time after a cell is infected but before it is capable of releasing virions. The duration of the eclipse, or the subsequent infectious, phase is non-exponential, but composed of stages. We derive the probability distribution of the reproduction number for these viral dynamics models, and show it is a negative binomial distribution in the case of constant viral release from infectious cells, and under the assumption of an excess of target cells. In a deterministic model, the ultimate in-host establishment or extinction of the viral infection depends entirely on whether the mean reproduction number is greater than, or less than, one, respectively. Here, the probability of extinction is determined by the probability distribution of R, not simply its mean value. In particular, we show that in some cases the probability of infection is not an increasing function of the mean reproduction number.

59 BASIC BIOLOGICAL SCIENCES↗

Seismic Contingency Auto Generator

This code takes in premade earthquake scenario XML files from USGS, power grid data, and converts them into a contingency file (.con file) that can be used by power grid solvers. Within the .con file are a number (Specified by the user) of contingencies that have randomly failed power transformers based on their likelihood of failure and peak ground acceleration (PGA) value around the transformer. The transformers' likelihood of failure was calculated based on a variety of finite element modeling on various transformer designed for specific transformer voltage classes. Parameters from these FEM were used to create generic fragility curves for transformers within a specific voltage class, which correspond with earthquake PGA values to produced a probability of failure for a given earthquake scenario. More refined versions of this process, such as specifying specific transformer design categories within a voltage class, could also be applied in future iterations of the software.

Vaagensmith, Bjorn [Idaho National Laboratory (INL↗

Deterministic Linear Time for Maximal Poisson‐Disk Sampling using Chocks without Rejection or Approximation

Abstract We show how to sample uniformly within the three‐sided region bounded by a circle, a radial ray, and a tangent, called a “chock.” By dividing a 2D planar rectangle into a background grid, and subtracting Poisson disks from grid squares, we are able to represent the available region for samples exactly using triangles and chocks. Uniform random samples are generated from chock areas precisely without rejection sampling. This provides the first implemented algorithm for precise maximal Poisson‐disk sampling in deterministic linear time. We prove O(n · M(b) log b), where n is the number of samples, b is the bits of numerical precision and M is the cost of multiplication. Prior methods have higher time complexity, take expected time, are non‐maximal, and/or are not Poisson‐disk distributions in the most precise mathematical sense. We fill this theoretical lacuna.

Mitchell, Scott A.↗

Ch3MS-RF: a random forest model for chemical characterization and improved quantification of unidentified atmospheric organics detected by chromatography–mass spectrometry techniques

Abstract. The chemical composition of ambient organic aerosols plays a critical role in driving their climate and health-relevant properties and holds important clues to the sources and formation mechanisms of secondary aerosol material. In most ambient atmospheric environments, this composition remains incompletely characterized, with the number of identifiable species consistently outnumbered by those that have no mass spectral matches in the literature or the National Institute of Standards and Technology/National Institutes of Health/Environmental Protection Agency (NIST/NIH/EPA) mass spectral databases, making them nearly impossible to definitively identify. This creates significant challenges in utilizing the full analytical capabilities of techniques which separate and generate spectra for complex environmental samples. In this work, we develop the use of machine learning techniques to quantify and characterize novel, or unidentifiable, organic material. This work introduces Ch3MS-RF (Chemical Characterization by Chromatography–Mass Spectrometry Random Forest Modeling), an open-source, R-based software tool, for efficient machine-learning-enabled characterization of compounds separated in chromatography–mass spectrometry applications but not identifiable by comparison to mass spectral databases. A random forest model is trained and tested on a known 130 component representative external standard to predict the response factors of novel environmental organics based on position in volatility–polarity space and mass spectrum, enabling the reproducible, efficient, and optimized quantification of novel environmental species. Quantification accuracy on a reserved 20 % test set randomly split from the external standard compound list indicates that random forest modeling significantly outperforms the commonly used methods in both precision and accuracy, with a median response factor percent error of −2 %, for modeled response factors, compared to > 15 %, for typically used proxy assignment-based methods. Chemical properties modeling, evaluated on the same reserved 20 % test set and an extrapolation set of species identified in ambient organic aerosol samples collected in the Amazon rainforest, also demonstrate robust performance. Extrapolation set property prediction mean absolute errors for carbon number, oxygen to carbon ratio (O : C), average carbon oxidation state (OSc‾), and vapor pressure are 1.8, 0.15, 0.25, and 1.0 (log(atm)), respectively. Extrapolation set out-of-sample R2 for all properties modeled are above 0.75, with the exception of vapor pressure. While predictive performance for vapor pressure is less robust compared to the other chemical properties modeled, random-forest-based modeling was significantly more accurate than other commonly used methods of vapor pressure prediction, decreasing the mean vapor pressure prediction error to 0.24 (log(atm)) from 0.55 (log(atm)) (chromatography-based vapor pressure prediction) and 1.2 (log(atm)) (chemical formula-based vapor pressure prediction). The random forest model significantly advances an untargeted analysis of the full scope of chemical speciation yielded by two-dimensional gas chromatography (GCxGC-MS) techniques and can be applied to gas chromatography coupled with electron ionization mass spectrometry (GC-MS) as well. It enables the accurate estimation of key chemical properties commonly utilized in the atmospheric chemistry community, which may be used to more efficiently identify important tracers for further individual analysis and to characterize compound populations uniquely formed under specific ambient conditions.

54 ENVIRONMENTAL SCIENCES↗

Order Parameter Engineering for Random Systems

The chemical short-range order (CSRO) in crystalline materials influences the properties, and its effect is significant in the context of multicomponent materials. Here, we propose a scheme for the CSRO parameter or Δ-parameter in terms of the number of like and unlike bonds in the multicomponent systems. The OPERA or Order Parameter Engineering for RAndom Systems scheme for semi-canonical and canonical ensembles is proposed. The proposed framework of Δ-parameter with OPERA framework can generate the single-phase supercell with desired CSRO without explicit energy calculations and provides a computationally efficient scheme for exploration of the CSRO. We demonstrate the applicability of the Δ-parameter as a scalar quantity for describing the CSRO in multicomponent alloys and oxides (FCC-CoCrNi, BCC-MoNbTaW, and (CoCuMgNiZn)O).

36 MATERIALS SCIENCE↗

Void formation in operator growth, entanglement, and unitarity

The structure of the Heisenberg evolution of operators plays a key role in explaining diverse processes in quantum many-body systems. In this paper, we discuss a new universal feature of operator evolution: an operator can develop a void during its evolution, where its nontrivial parts become separated by a region of identity operators. Such processes are present in both integrable and chaotic systems, and are required by unitarity. We show that void formation has important implications for unitarity of entanglement growth and generation of mutual information and multipartite entanglement. We study explicitly the probability distributions of void formation in a number of unitary circuit models, and conjecture that in a quantum chaotic system the distribution is given by the one we find in random unitary circuits, which we refer to as the random void distribution. We also show that random unitary circuits lead to the same pattern of entanglement growth for multiple intervals as in (1 + 1)-dimensional holographic CFTs after a global quench, which can be used to argue that the random void distribution leads to maximal entanglement growth.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Simulation of the Formation of DNA Double Strand Breaks and Chromosome Aberrations in Irradiated Cells

The formation of DNA double-strand breaks (DSBs) and chromosome aberrations is an important consequence of ionizing radiation. To simulate DNA double-strand breaks and the formation of chromosome aberrations, we have recently merged the codes RITRACKS (Relativistic Ion Tracks) and NASARTI (NASA Radiation Track Image). The program RITRACKS is a stochastic code developed to simulate detailed event-by-event radiation track structure: [1] This code is used to calculate the dose in voxels of 20 nm, in a volume containing simulated chromosomes, [2] The number of tracks in the volume is calculated for each simulation by sampling a Poisson distribution, with the distribution parameter obtained from the irradiation dose, ion type and energy. The program NASARTI generates the chromosomes present in a cell nucleus by random walks of 20 nm, corresponding to the size of the dose voxels, [3] The generated chromosomes are located within domains which may intertwine, and [4] Each segment of the random walks corresponds to approx. 2,000 DNA base pairs. NASARTI uses pre-calculated dose at each voxel to calculate the probability of DNA damage at each random walk segment. Using the location of double-strand breaks, possible rejoining between damaged segments is evaluated. This yields various types of chromosomes aberrations, including deletions, inversions, exchanges, etc. By performing the calculations using various types of radiations, it will be possible to obtain relative biological effectiveness (RBE) values for several types of chromosome aberrations.

Plante, Ianik↗