Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel projection algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

82 records · Page 5

Low synchronization Gram–Schmidt and generalized minimal residual algorithms

The Gram–Schmidt process uses orthogonal projection to construct the A = QR factorization of a matrix. When Q has linearly independent columns, the operator P = I - Q(QTQ)-1QT defines an orthogonal projection onto Q⊥. In finite precision, Q loses orthogonality as the factorization progresses. A family of approximate projections is derived with the form P = I - QTQT, with correction matrix T. When T = (QTQ)-1, and T is triangular, it is postulated that the best achievable orthogonality is $\mathcal{O}(ε)\mathcal{K}(A)$. We present new variants of modified (MGS) and classical Gram–Schmidt algorithms that require one global reduction step. An interesting form of the projector leads to a compact WY representation for MGS. In particular, the inverse compact WY MGS algorithm is equivalent to a lower triangular solve. Our main contribution is to introduce a backward normalization lag into the compact WY representation, resulting in a $\mathcal{O}(ε)\mathcal{K}[r_0, AV_m])$ stable Generalized Minimal Residual Method (GMRES) algorithm that requires only one global reduce per iteration. Finally, further improvements in performance are achieved by accelerating GMRES on GPUs.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Simulation of Lindblad's Equation

Constructing fast quantum logic gates is critical to building a scalable quantum computer. We consider a qudit, a quantum version of a bit that can take an arbitrary number of states, coupled with a cavity. In this project, we wish to force the qudit to reach the 0-state, for any possible initial state. The coupled system changes in time according to Lindblad’s equation, an ordinary differential equation on the density matrix of the quantum system. Lindblad’s equation contains some parameters that we can control, so-called control functions. We seek control functions which force the qudit to the 0-state within 2 microseconds, which is much faster than what is currently done in practice. The search method is gradient descent, a numerical optimization method that uses gradient information to iteratively improve the control parameters. My contribution to this project is an attempt to speed up the computation of the gradient. It currently takes about 40 seconds to compute the gradient which involves solving a set of ODEs sequentially. Current supercomputers have thousands of cores, but sequential computations can only make use of 1 core at a time. We wish to divide up the work better, so that we can use many more cores at once. To this end, we have implemented the Multigrid Reduction in Time (MGRIT) algorithm. We perform a systematic parameter search on how to best apply this algorithm. Results indicate a 25 percent speed up for solving Lindblad’s equation and determining how close the final state is the 0-state.

97 MATHEMATICS AND COMPUTING↗

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits↗

The Kokkos Ecosystem [Brief]

In 2016/2017, the field of High-Performance Computing (HPC) entered a new era driven by fundamental physics challenges to produce ever more energy and cost-efficient processors. Since the convergence on the Message-Passing Interface (MPI) standard in the mid-1990s, application developers enjoyed a seemingly static view of the underlying machine — that of a distributed collection of homogeneous nodes executing in collaboration. However, after almost two decades of dominance, the sole use of MPI to derive parallelism acted as a limiter to improved future performance. While MPI is widely expected to continue to function as the basic mechanism for communication between compute nodes for the immediate future, additional parallelism is required on the computing node itself if high performance and efficiency goals are to be realized. When reviewing the architectures of the top HPC systems today, the change in paradigm is clear: the compute nodes of the leading machines in the world are either powered by many-core chips with a few dozen cores each, or use heterogeneous designs, where traditional CPUs marshal work to massively parallel compute accelerators which has as many as 200,000 processing threads in flight simultaneously. Complicating matters further for application developers, each processor vendor has its own preferred way of writing code for their architecture.The Kokkos EcoSystem was released by Sandia in 2017 to address this new era in HPC system design by providing a vendor independent performance portable programming system for scientific, engineering, and mathematical software applications written in the C++ programming language. Using Kokkos, application developers can be more productive because they will not have to create and maintain separate versions of their software for each architecture, nor will they have to be experts in each architecture's peculiar requirements. Instead, they will have a single method of programming for the diverse set of modern HPC architectures. While Kokkos started in 2011 as a programming model only, it soon became clear that complex applications needed more. It is also critical to have a portable mathematical functions and developers need tools to debug their applications, gain insight into the performance characteristics of their codes and tune algorithm performance parameters through automated processes. The Kokkos EcoSystem addresses those needs through its three main components: the Kokkos Core programming model, the Kokkos Kernels math library, and the Kokkos Tools project.

97 MATHEMATICS AND COMPUTING↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Visualization at exascale: Making it all work with VTK-m

The VTK-m software library enables scientific visualization on exascale-class supercomputers. Exascale machines are particularly challenging for software development in part because they use GPU accelerators to provide the vast majority of their computational throughput. Algorithmic designs for GPUs and GPU-centric computing often deviate from those that worked well on previous generations of high-performance computers that relied on traditional CPUs. Fortunately, VTK-m provides scientific visualization algorithms for GPUs and other accelerators. VTK-m also provides a framework that simplifies the implementation of new algorithms and adds a porting layer to work across multiple processor types. This paper describes the main challenges encountered when making scientific visualization available at exascale. Here, we document the surprises and obstacles faced when moving from pre-exascale platforms to the final exascale designs and the performance on those systems including scaling studies on Frontier, an exascale machine with over 37,000 AMD GPUs. We also report on the integration of VTK-m with other exascale software technologies. Finally, we show how VTK-m helps scientific discovery for applications such as fusion and particle acceleration that leverage an exascale supercomputer.

97 MATHEMATICS AND COMPUTING↗

HIPPO – A Software Platform for Electricity Market Research and Development

The goal of this project is to provide Regional transmission organizations (RTOs) and independent system operators (ISOs) a market design and prototyping software, High-Performance Power-Grid Optimization (HIPPO), that they can evaluate electricity market design options, calculate market planning strategies and operational performance. With the high standards and strict reliability requirements for operating power systems, impacts of new technologies need to be fully investigated prior to any consideration for adoption. A market design and prototyping software tool which can be used to prototype electricity market design options, to calculate market planning strategies and operational performance with high precision, and to investigate the impacts for integrating future power grid technologies will be valuable to RTOs/ISOs who operate power systems, to vendors like GE and ABB who provide the market solvers, and to market participants and researchers who are actively doing market research. HIPPO is a such tool that can be used to improve the current market operations and provide capabilities for rigorous forward-looking design and prototyping of next-generation energy markets. HIPPO has a high-resolution model for the day-ahead SCUC, which was validated with MISO and GE-Grid Solutions. HIPPO is built with parallel and distributed computing capabilities and can be executed in both multi-thread and high-performance computing (HPC) settings. This capability provides fast solution speed necessary to handle the larger and more complex SCUC problems of real-world cases and the potentially growing size and complexity of future scenarios. In addition, HIPPO has a concurrent optimizer (CO) which manages multiple algorithm executions simultaneously and leverages the advantages from different algorithms. This structure provides flexibility to better benchmark competing approaches. Highly accurate market model, fast solution technologies and flexible model and algorithm control are the features which will make HIPPO an extensible platform for developing and testing multiple approaches to meet a wide range of future market needs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗

Hybrid Solar System (Final Scientific/Technical Report)

GTI Energy (GTI) teamed with the University of California at Merced (UCM) to scaleup the hybrid solar system (HSS) technology for demonstrating its performance at the US Gypsum (USG) plant in Plaster City, California. The technology integrates two-stage concentrating solar collector with matching particle thermal transport and storage (TSS) system to deliver cost-effective, and on-demand distributed high temperature industrial process heat up to 600°C with solar thermal, in this case to a gypsum kettle, to reduce its fuel use and carbon footprint. Current solar technologies, which reach these temperatures, are not distributable (towers) or cost-effective (dish). The research team developed a conceptual system design for host site retrofit, including preliminary heat balance, process flow diagram, particle to process heat exchanger and equipment placements at the site. Subsequently, parallel efforts were carried out at UCM to design, build and test a 12 m long commercial scale prototype concentrating thermal-only collector system and at GTI to design, build and test a matching 650°C capable particle TTS system. The nominal 50 kWth collector consists of a parabolic trough and three 4 m long two-stage receivers in series. Prior to on-sun testing, a 4 m long receiver was fabricated and successfully tested at 650 °C in a laboratory setting for 100 hrs of continuous operation showing less than 15% radiation loss. A 7 m wide x 17 m long parabolic trough was then installed at UCM for on-sun testing of the 12 m long receiver, and concurrently several 4 m long receivers were built. The optics of the parabolic trough were calibrated, and on-sun test were carried out on 12 m long receivers. During tests, the intense solar radiation (53x) caused the absorber tubes in the receivers to bend, reducing the overall optical efficiency. To address the bending issue, a self-consistent algorithm that includes ray tracing, thermal and deformation models was developed to perform thermal stress analysis on absorbers for parabolic solar collectors. Results obtained with this algorithm showed a dramatic rise in deformation as absorber tube length increases. A combined efficiency parameter that includes the occluded area for the mounts was developed to obtain an optimized tube length obtained. Based on the results, a length of 2.7 m for the absorber + 0.2 m for the coupler was chosen to minimize any bending and optimize optical efficiency while maintaining ease of mounting. The associated particle TTS system was designed, built and successfully tested at GTI. It includes storage, receiving and lock hoppers and piping that simulates the transfer of captured solar energy to an actual industrial furnace. Tests over 77 charge-discharge cycles demonstrated <2% particle degradation, with no problematic particle accumulations and no flow interruptions. The piping pressure drop was about 5 psi. The team also worked with Stanley Consultants (Stanley) to prepare conceptual and preliminary engineering packages to facilitate follow-on development and commercialization efforts. These include process and instrumentation diagram’s (P&ID’s), general arrangements, electrical one-line, project definitions document, equipment data sheets, schedule, and construction cost estimate for 2 MWth system. Updated HSS technology commercialization and customer engagement plans and detailed costs and evaluated market trade-offs and manufacturing.

03 NATURAL GAS↗

Development of NDE/NDT Tools for High-Volume & High-Speed Inspection of CFRP Structures in Automotive Manufacturing

Main advantages of the air-coupled ultrasound testing (ACUT) and electromagnetic testing (EMT) techniques for NDE of CFRP composites were non-contact sensing, scalability for high-speed inspection, cost-effectiveness, and non-hazardous operation. Despite these advantages, no systems that would satisfy the project requirements were commercially available. Hence, one of the major efforts of the Michigan State University (MSU) team at the initial stage of the project was to close this technological gap by developing, optimizing, and validating array sensors that would provide sufficient sensitivity, spatial coverage, and resolution for robust defect detection. Optimization of the ACUT and EMT sensor designs was performed using experimentally validated finite element models. Initial experiments using array probes were conducted on relatively flat CFRP samples. In parallel, the MSU team designed and assembled a portable platform with two robotic arms. The robots were equipped with newly designed sensors that enabled high-speed NDE of curved CFRP parts. Presently, the developed robotic platform can be used as a demo/template NDE system, which is easily adaptable to manufacturing environments and in-line NDE. The ACUT NDE system developed by the MSU team used a high-power 4-channel pulser receiver for parallel data acquisition. The array probes were designed by stacking commercially available ACUT transducers, which operated in the frequency range between 100 kHz and 500 kHz. MSU optimized the excitation procedure and developed wave focusing cones so as to reduce the crosstalk between the transducers and to provide higher pulse repletion frequency (PRF). The through-transmission (TT) and single-side access (SSA) inspection modes were successfully implemented. In the TT-ACUT, structural defects in CFRP were detected by passing ultrasonic waves through the test part. Hence, the ACUT transmitters and receivers needed to be placed on the opposite sides of the test part. In the SSA-ACUT, guided waves (GW) were excited in the test part using the transmitters and were sensed by the receivers from the same side. Multi-channel TT-ACUT and SSA-ACUT provided high-speed NDE, and were successfully validated on CFRP test samples with interlaminar delaminations and other embedded defects The EM techniques developed by the MSU team included: 1) eddy current testing (ECT), 2) capacitive imaging (CI) and hybrid dual-mode imaging. In ECT, structural damage was detected in CFRP using coils sensor arrays. In ECT, the excitation magnetic field is generated by passing an alternating current through a coil, which is placed above the test sample. The excitation field penetrates the conductive sample and induces the eddy currents in its transect. In turn, the eddy currents generate the reaction field, which affects the total field sensed by a coil. Hence, the presence of structural flaws will alter the eddy current flow and the picked-up signal. ECT is mostly sensitive to local changes of the electric conductivity of the test sample, and CFRPs are mostly conductive in the direction of carbon fibers. Hence, ECT was well suited for the detection of fiber damage/fiber irregularities. The MSU team developed printed circuit boards (PCB) with coil sensor arrays optimized for NDE of CFRP. Unlike most commercial probes designed for ECT of metallic structures, the MSU array probes were designed for operation in [1-10] MHz frequency range, which was optimal for low-conductive CFRP. Multiple sensing topologies (coil groups excitation/sensing arrangements) were implemented and successfully validated. Capacitive Imaging (CI) technique developed by MSU was complementary to ECT. In contrast to ECT, which was sensitive to local changes of the electrical conductivity, the CI was sensitive to local changes of the dielectric constant. Therefore, CI could provide information about matrix damage/matrix irregularities in CFRP. The MSU CI sensor arrays were made of multiple circular or rectangular open-plate capacitors printed on PCB. Sensors of this type are not commercially available. In addition to ECT and CI, the MSU team developed a hybrid (dual-mode) inductive/capacitive measurement technique that synergistically combined the benefits of inductive and capacitive sensing for rapid NDE of fiber reinforced polymer (FRP) composite structures. Fiber damage and fiber irregularities in FRPs were detected by configuring hybrid sensors as coil sensors. Similarly, matrix damage, matrix irregularities and interlaminar delaminations were detected by configuring hybrid sensors as capacitive sensors. ECT and CI were performed sequentially by means of electronic switching. Hence, eliminating the need for mounting two separate sensor arrays on the probe. Portable robotic platform was developed by MSU for multi-technique high-speed NDE of CFRP test parts. The platform had two 6-axis robots, which enabled inspection of curved parts in approximately a 6×6×6 ft 3 active scan area. On the software side, the MSU team integrated scripts for NDE hardware control with scripts for robot motion control. MSU also implemented automated path planning for the robots, reconstruction of part’s surfaces via stereovision, 3D rendering of inspection data, and image processing algorithms for enhanced defect detection. Automotive composite parts manufactured by Plasan Composites from Phase I were used to validate the ACUT and EMT techniques on representative testbeds. Among those parts were three X-braces for a Dodge Viper, one composite calibration plaque with known defects at known locations, and four other test sections, including sections from a front splitter, a corner section from a composite hood, and a high-pressure RTM panel made using non crimp fabric. Other test samples included CFRP and GFRP calibration plates with fiber/matrix defects fabricated at MSU/CVRC.

36 MATERIALS SCIENCE↗