Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Graph processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

D-Side: A Facility and Workforce Planning Group Multi-criteria Decision Support System for Johnson Space Center

"To understand and protect our home planet, to explore the universe and search for life, and to inspire the next generation of explorers" is NASA's mission. The Systems Management Office at Johnson Space Center (JSC) is searching for methods to effectively manage the Center's resources to meet NASA's mission. D-Side is a group multi-criteria decision support system (GMDSS) developed to support facility decisions at JSC. D-Side uses a series of sequential and structured processes to plot facilities in a three-dimensional (3-D) graph on the basis of each facility alignment with NASA's mission and goals, the extent to which other facilities are dependent on the facility, and the dollar value of capital investments that have been postponed at the facility relative to the facility replacement value. A similarity factor rank orders facilities based on their Euclidean distance from Ideal and Nadir points. These similarity factors are then used to allocate capital improvement resources across facilities. We also present a parallel model that can be used to support decisions concerning allocation of human resources investments across workforce units. Finally, we present results from a pilot study where 12 experienced facility managers from NASA used D-Side and the organization's current approach to rank order and allocate funds for capital improvement across 20 facilities. Users evaluated D-Side favorably in terms of ease of use, the quality of the decision-making process, decision quality, and overall value-added. Their evaluations of D-Side were significantly more favorable than their evaluations of the current approach. Keywords: NASA, Multi-Criteria Decision Making, Decision Support System, AHP, Euclidean Distance, 3-D Modeling, Facility Planning, Workforce Planning.

Tavana, Madjid↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

On the Feasibility of Using Reduced-Precision Tensor Core Operations for Graph Analytics

Today’s data-driven analytics and machine learning workload have been largely driven by the General-PurposeGraphics Processing Units (GPGPUs). To accelerate dense matrix multiplications on the GPUs, Tensor Core Units (TCUs) have been introduced in recent years. In this paper, we study linear-algebra-based and vertex-centric algorithms for various graph kernels on the GPUs with an objective of applying this new hardware feature to graph applications. We identify the potential stages in these graph kernels that can be executed on the Tensor Core Units. In particular, we leverage the reformulation of the reduction and scan operations in terms of matrix multiplication [1]on the TCUs. We demonstrate that executing these operations on the TCUs, available inside different graph kernels, can assist in establishing an end-to-end pipeline on the GPGPUs without depending on hand-tuned external libraries and still can deliver comparable performance for various graph analytics.

Graph algorithms, GPU computing↗

Bounce-averaged theory in arbitrary multi-well plasmas: solution domains and the graph structure of their connections

Bounce-averaged theories provide a framework for simulating relatively slow processes, such as collisional transport and quasilinear diffusion, by averaging these processes over the fast periodic motions of a particle on a closed orbit. This procedure dramatically increases the characteristic time scale and reduces the dimensionality of the modelled system. The natural coordinates for such calculations are the constants of motion (COM) of the fast particle motion, which by definition do not change during an orbit. However, for sufficiently complicated fields – particularly in the presence of local maxima of the electric potential and magnetic field – the COM are not sufficient to specify the particle trajectory. In such cases, multiple domains in COM space must be used to solve the problem, with boundary conditions enforced between the domains to ensure continuity and particle conservation. Previously, these domains have been imposed by hand, or by recognising local maxima in the fields, limiting the flexibility of bounce-averaged simulations. Here, we present a general set of conditions for identifying consistent domains and the boundary condition connections between the domains, allowing the application of bounce-averaged theories in arbitrarily complicated and dynamically evolving electromagnetic field geometries. We also show how the connections between the domains can be represented by a directed graph, which can help to succinctly represent the trajectory bifurcation structure.

fusion plasma↗

LigninGraphs: lignin structure determination with multiscale graph modeling

Lignin is an aromatic biopolymer found in ubiquitous sources of woody biomass. Designing and optimizing lignin valorization processes requires a fundamental understanding of lignin structures. Experimental characterization techniques, such as 2D-heteronuclear single quantum coherence (HSQC) nuclear magnetic resonance (NMR) spectra, could elucidate the global properties of the polymer molecules. Computer models could extend the resolution of experiments by representing structures at the molecular and atomistic scales. We introduce a graph-based multiscale modeling framework for lignin structure generation and visualization. The framework employs accelerated rejection-free polymerization and hierarchical Metropolis Monte Carlo optimization algorithms. We obtain structure libraries for various lignin feedstocks based on literature and new experimental NMR data for poplar wood, pinewood, and herbaceous lignin. The framework could guide researchers towards feasible lignin structures, efficient space exploration, and future kinetics modeling. Its software implementation in Python, LigninGraphs, is open-source and available on GitHub.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An automated procedure built on MTEX for reconstructing deformation twin hierarchies from electron backscattered diffraction datasets of heavily twinned microstructures

Here we present a set of algorithms built on the MTEX and MATLAB graph toolboxes for automatic reconstruction of deformation twin hierarchies from Electron Backscatter Diffraction (EBSD) datasets with a focus on developing methods for heavily twinned microstructures (twin fractions >0.5). The algorithms address key issues arising at large strains, mainly: missing twin relationships, grouping of heavily deformed grain fragments into families of similar orientation originating from a single initial grain, identification of parent fragments for large twin volume fractions, and classification of families having twin relationships with multiple families. To facilitate the development of these algorithms, large-grained ultra-high purity α-Ti deformed in compression along two directions is investigated. Graphs are utilized to handle non-local geometric merging and to represent relationships throughout the reconstruction process. When determining if a grain fragment is from the undeformed microstructure, the combined metrics of the fragment's orientation volume fraction in the initial texture and the directed graph centrality measure of out-closeness (the number of nodes reached in a graph from a given node) are essential. To address automation in reconstructing the sequence of twinning and relating fragments originating from a single grain in the initial microstructure, the twin family tree is formulated as a minimum spanning tree emanating from the initial grain family. A scheme constructing the distances associated with twin relationship comprising the spanning tree is developed, and a novel quasi-directional Prim spanning tree algorithm is used to determine the twin family tree. The procedure is demonstrated to significantly improve the level of automation in reconstructing twin hierarchies in heavily twinned microstructure compared to other methodologies in literature. The procedure can readily be applied to analyses of twinning in metals, as well as provide an approach for routinely extracting twin statistics at larger deformation levels than previously possible. Significantly, the procedure is demonstrated to be capable of identifying third generation twinning in α-Ti microstructures.

36 MATERIALS SCIENCE↗

GraMeR: Gra ph Me ta R einforcement learning for multi-objective influence maximization

Influence maximization (IM) is a combinatorial problem of identifying a subset of seed nodes in a network (graph), which when activated, provide a maximal spread of influence in the network for a given diffusion model and a budget for seed set size. IM has numerous applications such as viral marketing, epidemic control, sensor placement and other network-related tasks. However, its practical uses are limited due to the computational complexity of current algorithms. Recently, deep reinforcement learning has been leveraged to solve IM in order to ease the computational burden. However, there are serious limitations in current approaches, including narrow IM formulation that only consider influence via spread and ignore self-activation, low scalability to large graphs, and lack of generalizability across graph families leading to a large running time for every test network. In this work, we address these limitations through a unique approach that involves: (1) Formulating a generic IM problem as a Markov decision process that handles both intrinsic and influence activations; (2)incorporating generalizability via meta-learning across graph families. There are previous works that combine deep reinforcement learning with graph neural network, but this work solves a more realistic IM problem and incorporates generalizability across graphs via meta reinforcement learning. Extensive experiments are carried out in various standard networks to validate performance of the proposed Graph Meta Reinforcement learning (GraMeR) framework. Finally, the results indicate that GraMeR is multiple orders faster and generic than conventional approaches when applied on small to medium scale graphs.

97 MATHEMATICS AND COMPUTING↗

Optimal adjustment sets for causal query estimation in partially observed biomolecular networks

Abstract Causal query estimation in biomolecular networks commonly selects a ‘valid adjustment set’, i.e. a subset of network variables that eliminates the bias of the estimator. A same query may have multiple valid adjustment sets, each with a different variance. When networks are partially observed, current methods use graph-based criteria to find an adjustment set that minimizes asymptotic variance. Unfortunately, many models that share the same graph topology, and therefore same functional dependencies, may differ in the processes that generate the observational data. In these cases, the topology-based criteria fail to distinguish the variances of the adjustment sets. This deficiency can lead to sub-optimal adjustment sets, and to miss-characterization of the effect of the intervention. We propose an approach for deriving ‘optimal adjustment sets’ that takes into account the nature of the data, bias and finite-sample variance of the estimator, and cost. It empirically learns the data generating processes from historical experimental data, and characterizes the properties of the estimators by simulation. We demonstrate the utility of the proposed approach in four biomolecular Case studies with different topologies and different data generation processes. The implementation and reproducible Case studies are at https://github.com/srtaheri/OptimalAdjustmentSet.

59 BASIC BIOLOGICAL SCIENCES↗

A Performance and Energy Study of GPU-Resident Preconditioners for Conjugate Gradient Solvers: In the Context of Existing and Novel Approaches

Optimizing a particular subprogram out of the set of Basic (sparse) Linear Algebra Subprograms (BLAS) for a given architecture is a common topic of research. In applications, however, these BLAS functions rarely appear in isolation; usually, many of them are used together, in various combinations and with varying inputs. As the need to solve a large, sparse linear system is ubiquitous throughout HPC applications, linear solvers constitute a realistic, sufficiently complex and well-defined representative use case for composite BLAS routines. To this end, based on a representative set of matrices drawn from a diverse set of fields, we present a framework to study, from the performance and energy perspective, the efficacy of GPU- resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process. The development of this preconditioner was motivated by solving large graph Laplacian linear systems, for which the existing preconditioners either perform slow on GPU-based platforms or are not applicable. We compare the performance of these preconditioners on different hardware accelerator architectures, i.e., AMD MI250X, MI100, Nvidia A100, V100, and Jetson. Our experiments reveal performance trade-offs and provide information on how to select the best strategy for the given linear system, dictated by its properties, and the platform of interest. We demonstrate the application of our novel preconditioner for solving CG and graph Laplacian systems. Overall, the framework can be utilized as a benchmark to guide informed decisions in choosing a specific preconditioner, i.e., whether it is better to rely on the performance of a triangular solver or on the performance of sparse matrix-vector product. Finally, by considering power consumption to solve the linear systems, we report the energy footprint for the solvers.

Preconditioned Conjugate Gradient, GPUs, iterative↗

Detailed study on acceleration and propagation of energetic protons and electrons in the magnetotail during substorm activity

High time resolution measurements of energetic particles and magnetic field measurements by the IMP 8 satellite in the distant magnetotail are presented for November 26, 1973, when exceptionally intense particle bursts were detected by both the IMP 7 and 8 spacecraft. During the onset of the most intense burst as well as at other times, oppositely directed anisotropies of protons and electrons parallel to the tail field and lasting up to about 60 sec were observed, implying the presence of field-aligned electric fields. The particle and field observations are discussed in the context of proposed mechanisms for the acceleration of particles during various dynamical magnetospheric processes. Satellite instrument readings are presented through the extensive use of graphs.

Kirsch, E.↗

Domain decomposition methods in aerodynamics

Compressible Euler equations are solved for two-dimensional problems by a preconditioned conjugate gradient-like technique. An approximate Riemann solver is used to compute the numerical fluxes to second order accuracy in space. Two ways to achieve parallelism are tested, one which makes use of parallelism inherent in triangular solves and the other which employs domain decomposition techniques. The vectorization/parallelism in triangular solves is realized by the use of a recording technique called wavefront ordering. This process involves the interpretation of the triangular matrix as a directed graph and the analysis of the data dependencies. It is noted that the factorization can also be done in parallel with the wave front ordering. The performances of two ways of partitioning the domain, strips and slabs, are compared. Results on Cray YMP are reported for an inviscid transonic test case. The performances of linear algebra kernels are also reported.

Venkatakrishnan, V.↗

A comparison of multiprocessor scheduling methods for iterative data flow architectures

A comparative study is made between the Algorithm to Architecture Mapping Model (ATAMM) and three other related multiprocessing models from the published literature. The primary focus of all four models is the non-preemptive scheduling of large-grain iterative data flow graphs as required in real-time systems, control applications, signal processing, and pipelined computations. Important characteristics of the models such as injection control, dynamic assignment, multiple node instantiations, static optimum unfolding, range-chart guided scheduling, and mathematical optimization are identified. The models from the literature are compared with the ATAMM for performance, scheduling methods, memory requirements, and complexity of scheduling and design procedures.

Storch, Matthew↗

Data-flow parallelism for high-energy and nuclear physics computing frameworks

The processing tasks of a scientific workflow in high-energy and nuclear physics (HENP) can typically be represented as a directed acyclic graph formed according to the data flow—i.e. the data dependencies among algorithms executed as part of the workflow. With this representation, an HENP computing framework can optimally execute a workflow, exploiting the parallelism inherent among independent tasks. Despite such a natural description of a workflow, most HENP frameworks do not make use of technologies that provide concurrent execution of graph-based tasking structures. In this session, we describe Fermilab efforts to adopt a graph-based technology (specifically Intel’s oneTBB flow graph) for meeting the framework needs of its experiments, notably DUNE. After introducing the physics DUNE intends to explore, we will show that all common processing idioms supported by current HENP frameworks can naturally be supported by oneTBB’s data-flow technology, optimally leveraging the concurrent capabilities of the machine. In addition, we discuss collaborative efforts between Fermilab and the Intel oneTBB development team, who is considering improvements to the flow-graph technology to better support HENP use cases.

43 PARTICLE ACCELERATORS↗

Natural Language Processing-Enhanced Nuclear Industry Operating Experience Data Analysis: Aggregation and Interpretation of Multi-Report Analysis Results

Industry-wide operating experience is a critical source of raw data for reliability and risk model parameter estimations for nuclear power plants. A large portion of operating experience data are failure events stored as reports that contain unstructured data, such as narratives. In current practice, a failure report is usually reviewed and manually coded by analysts. The coding is based on extracting several event characteristics such as system name, component type, sub-part type, failure mode, and failure cause. Event narratives are mostly used to help understand events and extract their characteristics. In this line of research, we aim to maximize the usage of event narratives by leveraging natural language processing (NLP) methods to automatically convert an event narrative to a causal graph. This research has promise to improve physical understanding of failure initiation and propagation and to facilitate use of non-failure data (e.g., near-misses and degradations) to complement the limited data pool of failures. In our previous work, we developed an NLP tool and applied it to analyze a number of licensee event reports submitted by U.S. nuclear power plants to the Nuclear Regulatory Commission. In this paper, we will report our recent research progress in aggregating the results of multiple reports, developing network model(s), and drawing statistical insights.

99 GENERAL AND MISCELLANEOUS↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

A Multilevel Approach For SolvingLarge-Scale QUBO Problems With Noisy Hybrid Quantum Approximate Optimization

Quantum approximate optimization is one ofthe promising candidates for useful quantum computation,particularly in the context of finding approximate solutionsto Quadratic Unconstrained Binary Optimization (QUBO)problems. However, the existing quantum processing units(QPUs) are of relatively small size, and canonical mappingsof QUBO via the Ising model require one qubit per vari-able, rendering direct large-scale optimization infeasible.In classical optimization, a general strategy for addressingmany large-scale problems is via multilevel/multigrid meth-ods, where the large target problem is iteratively coarsenedand the global solution is constructed from multiple small-scale optimization runs. In this work, we experimentallytest how existing QPUs perform when used as a sub-solverwithin such a multilevel strategy. To this aim, we com-bine and extend (via additional classical processing steps)the recently proposed Noise-Directed Adaptive Remapping(NDAR) and Quantum Relax&Round (QRR) algorithms.We first demonstrate the effectiveness of our heuristicextensions on Rigetti’s superconducting transmon deviceAnkaa-2. We find approximate solutions to10instances offully connected82-qubit Sherrington-Kirkpatrick graphswith random integer-valued coefficients obtaining normal-ized approximation ratios (ARs) in the range∼0.98−1.0,and the same class with real-valued coefficients (ARs∼0.94−1.0). Then, we implement the extended NDAR andQRR algorithms as subsolvers in the multilevel algorithmfor6large-scale graphs with at most∼27,000variables.In practice, the QPU (with classical post-processing steps)is used to find approximate solutions to dozens of at most82-qubit problems, which are iteratively used to constructthe global solution. We observe that quantum optimizationresults are competitive in terms of the quality of solutionswhen compared to classical heuristics used as subsolverswithin the multilevel approach.Reproducibility: source code and data are available at[TBA upon acceptance]

quantum computing↗

Off-line programming motion and process commands for robotic welding of Space Shuttle main engines

The off-line-programming software and hardware being developed for robotic welding of the Space Shuttle main engine are described and illustrated with diagrams, drawings, graphs, and photographs. The menu-driven workstation-based interactive programming system is designed to permit generation of both motion and process commands for the robotic workcell by weld engineers (with only limited knowledge of programming or CAD systems) on the production floor. Consideration is given to the user interface, geometric-sources interfaces, overall menu structure, weld-parameter data base, and displays of run time and archived data. Ongoing efforts to address limitations related to automatic-downhand-configuration coordinated motion, a lack of source codes for the motion-control software, CAD data incompatibility, interfacing with the robotic workcell, and definition of the welding data base are discussed.

Ruokangas, C. C.↗

Fading of Jupiter's South Equatorial Belt

One of Jupiter's most dominant features, the South Equatorial Belt, has historically gone through a "fading" cycle. The usual dark, brownish clouds turn white, and after a period of time, the region returns to its normal color. Understanding this phenomenon, the latest occurring in 2010, will increase our knowledge of planetary atmospheres. Using the near infrared camera, NSFCAM2, at NASA's Infrared Telescope Facility in Hawaii, images were taken of Jupiter accompanied by data describing the circumstances of each observation. These images are then processed and reduced through an IDL program. By scanning the central meridian of the planet, graphs were produced plotting the average values across the central meridian, which are used to find variations in the region of interest. Calculations using Albert4, a FORTRAN program that calculates the upwelling reflected sunlight from a designated cloud model, can be used to determine the effects of a model atmosphere due to various absorption, scattering, and emission processes. Spectra that were produced show ammonia bands in the South Equatorial Belt. So far, we can deduce from this information that an upwelling of ammonia particles caused a cloud layer to cover up the region. Further investigations using Albert4 and other models will help us to constrain better the chemical make up of the cloud and its location in the atmosphere.

near-infrared↗