Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Architecture patterns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Experimental investigation on the effects of fabric architectures on mechanical and damage behaviors of carbon/epoxy woven composites

The mechanical behaviors and damage evolutions of carbon/epoxy woven fabric composites with three different geometries, i.e., one plain weave and two twill weave patterns with different areal densities, are studied under tensile loading. The effects of weave patterns on mechanical properties are investigated by monotonic and cyclic tension tests. Remarkable variations in stress–strain curve, Poisson’s ratio, residual strain and strain map exist in the three composites. Crimp ratio is found to be a critical factor to govern the mechanical properties. With smaller crimp ratio, a quasi-linear stress–strain curve with higher elastic modulus and strength is observed. The stress–strain curves of composites with higher crimp ratio contain transition stages with significant tangent modulus degradation. Elastic modulus, strength and damage initiation are all correlated with the crimp ratio linearly regardless of the fabric pattern. Dramatic nonlinear evolution in Poisson’s ratio occurs in the composite with higher crimp ratio. Cyclic tension results indicate that the residual strain is a more appropriate damage indicator than the unloading elastic modulus. Microstructure examination shows that damage developments are essentially related to the fabric geometry, and result in various mechanical behaviors. This work provides important insights into the geometry-deformation mechanism-mechanical property relationship of the woven composites.

36 MATERIALS SCIENCE↗

Integrated System Planning: Emerging Software Requirements in the Power Industry

Power system planning software remains fragmented across organizational boundaries, with specialized tools for capacity expansion, production cost modeling, power flow, and dynamic analysis operating on incompatible data models and assumptions. This article argues that the fragmentation is not merely a technical problem but a predictable consequence of Conway's law: software architectures mirror the departmental structures within which they are developed. Regulatory milestones like Federal Energy Regulatory Commission (FERC) Order 888 formalized these divisions, but the roots trace back to the distinct engineering disciplines-mechanical, chemical, and electrical-that staffed generation and transmission planning departments in vertically integrated utilities. As the industry moves toward integrated system planning (ISP) that coordinates generation, transmission, and distribution investment decisions, the software ecosystem must evolve accordingly. We identify five categories of software requirements to enable this transition: coherent data inputs decoupled from individual applications, unified and extensible data schemas, modular component representations that support multiple abstraction levels, lifecycle management of planning datasets, and well-defined application programming interface (API) contracts that separate data exchange from algorithmic control. We examine how these requirements interact with three common workflow patterns-serial gate clearing, sequential multiapplication, and convergence oriented-and discuss the interface design principles each demands. We then outline a vision for platform-based planning architectures where specialized analytical services compose through standardized interfaces and where artificial intelligence (AI)/machine learning (ML) tools augment decision support within a disciplined software infrastructure. The practices proposed here offer a path from today's siloed tool collections toward collaborative planning ecosystems capable of handling the complexity of modern power system transformation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An integrated Ka/Ku-band payload for personal, mobile and private business communications

The Canadian Department of Communications has been studying options for a government-sponsored demonstration payload to be launched before the end of the century. A summary of the proposed system concepts and network architectures for providing an advanced private business network service at Ku-band and personal and mobile communications at Ka-band is presented. The system aspects addressed include coverage patterns, traffic capacity, and grade of service, multiple access options as well as special problems, such as Doppler in mobile applications. Earth terminal types and the advanced payload concept proposed in a feasibility study for the demonstration mission are described. This concept is a combined Ka-band/Ku-band payload which incorporates a number of advanced satellite technologies including a group demodulator to convert single-channel-per-carrier frequency division multiple access uplink signals to a time division multiplex downlink, on-board signal regeneration, and baseband switching to support packet switched data operation. The on-board processing capability of the payload provides a hubless VSAT architecture which permits single-hop full mesh interconnectivity. The Ka-band and Ku-band portions of the payload are fully integrated through an on-board switch, thereby providing the capability for fully integrated services, such as using the Ku-band VSAT terminals as gateway stations for the Ka-band personal and mobile communications services.

Hayes, Edward J.↗

ClimateNet: an expert-labeled open dataset and deep learning architecture for enabling high-precision analyses of extreme weather

Abstract. Identifying, detecting, and localizing extreme weather events is a crucial first step in understanding how they may vary under different climate change scenarios. Pattern recognition tasks such as classification, object detection, and segmentation (i.e., pixel-level classification) have remained challenging problems in the weather and climate sciences. While there exist many empirical heuristics for detecting extreme events, the disparities between the output of these different methods even for a single event are large and often difficult to reconcile. Given the success of deep learning (DL) in tackling similar problems in computer vision, we advocate a DL-based approach. DL, however, works best in the context of supervised learning – when labeled datasets are readily available. Reliable labeled training data for extreme weather and climate events is scarce. We create “ClimateNet” – an open, community-sourced human-expert-labeled curated dataset that captures tropical cyclones (TCs) and atmospheric rivers (ARs) in high-resolution climate model output from a simulation of a recent historical period. We use the curated ClimateNet dataset to train a state-of-the-art DL model for pixel-level identification – i.e., segmentation – of TCs and ARs. We then apply the trained DL model to historical and climate change scenarios simulated by the Community Atmospheric Model (CAM5.1) and show that the DL model accurately segments the data into TCs, ARs, or “the background” at a pixel level. Further, we show how the segmentation results can be used to conduct spatially and temporally precise analytics by quantifying distributions of extreme precipitation conditioned on event types (TC or AR) at regional scales. The key contribution of this work is that it paves the way for DL-based automated, high-fidelity, and highly precise analytics of climate data using a curated expert-labeled dataset – ClimateNet. ClimateNet and the DL-based segmentation method provide several unique capabilities: (i) they can be used to calculate a variety of TC and AR statistics at a fine-grained level; (ii) they can be applied to different climate scenarios and different datasets without tuning as they do not rely on threshold conditions; and (iii) the proposed DL method is suitable for rapidly analyzing large amounts of climate model output. While our study has been conducted for two important extreme weather patterns (TCs and ARs) in simulation datasets, we believe that this methodology can be applied to a much broader class of patterns and applied to observational and reanalysis data products via transfer learning.

54 ENVIRONMENTAL SCIENCES↗

Ultrasonic Nondestructive Evaluation Techniques Applied to the Quantitative Characterization of Textile Composite Materials

In this Progress Report, we describe our further development of advanced ultrasonic nondestructive evaluation methods applied to the characterization of anisotropic materials. We present images obtained from experimental measurements of ultrasonic diffraction patterns transmitted through water only and transmitted through water and a thin woven composite. All images of diffraction patterns have been included on the accompanying CD-ROM in the JPEG format and Adobe TM Portable Document Format (PDF), in addition to the inclusion of hardcopies of the images contained in this report. In our previous semi-annual Progress Report (NAG 1-1848, December, 1996), we proposed a simple model to simulate the effect of a thin woven composite on an insonifying ultrasonic pressure field. This initial approach provided an avenue to begin development of a robust measurement method for nondestructive evaluation of anisotropic materials. In this Progress Report, we extend that work by performing experimental measurements on a single layer of a five-harness biaxial woven composite to investigate how a thin, yet architecturally complex, material interacts with the insonifying ultrasonic field. In Section 2 of this Progress Report we describe the experimental arrangement and methods for data acquisition of the ultrasonic diffraction patterns upon transmission through a thin woven composite. We also briefly describe the thin composite specimen investigated. Section 3 details the analysis of the experimental data followed by the experimental results in Section 4. Finally, a discussion of the observations and conclusions is found in Section 5.

Miller, James G.↗

Transactional Knowledge Graph Generation To Model Adversarial Activities

A Knowledge Graph (KG) is a formal and structured representation of facts, relationships, and semantic descriptions of a set of entities. Traditionally, KGs are used to describe metadata about entities and to provide additional context to target application results. Many real-world domains also involve temporal interactions between entities in addition to the metadata data. Modeling these attributed transactions is a critical requirement when using KGs in complex real-world applications. Modeling adversarial activities is one such application that develops methodology and tools to produce realistic large-scale background activity graphs that include embedded Weapons of Mass Destruction (WMD) activity patterns. We present a novel platform for constructing a transactional knowledge graph from a diverse set of sources. We present the core components and architecture of the framework, and a use case for generating a background knowledge graph and WMD activity template to evaluate network alignment and subgraph matching algorithms.

Purohit, Sumit↗

Baylor University Campus-Wide Deep Dive

In January 2020, staff members from the Engagement and Performance Operations Center (EPOC) and the Lonestar Education And Research Network (LEARN) met with researchers and staff at Baylor University for the purpose of a Campus-Wide Deep Dive into research drivers. The goal of this meeting was to help characterize the requirements for five campus research use cases and to enable cyberinfrastructure support staff to better understand the needs of the researchers they support. Profiled scientific use cases included: - Experimental High Energy Physics (HEP) - Proton Computed Tomography (pCT) - Nutrition and Relation to Digestive Microbiome - Baylor University Core Research Facilities - Molecular Quantum-dot Cellular Automata (QCA), and Material Science of Quantum Computing - Modeling and Simulation of Low-Dimensional and Nano-Structured Materials - Computational Fluid Dynamics Material for this event included the written documentation from each of the research areas at Baylor University, documentation about the current state of technology support, and a write-up of the discussion that took place in person. The Case Studies highlighted the ongoing challenges that Baylor University has in supporting a cross-section of established and emerging research use cases. Each Case Study mentioned unique challenges which were summarized into common needs. These included: - Tradeoffs for network/software security, and usability of the resulting infrastructure. Better communication to set expectations and understand realities is required. - Computation use on campus is widespread and healthy. While no major problems were uncovered, upgrades to maintain current usage patterns and encourage growth will be required. - Storage is a critical need for enterprise use cases and research. In particular, a campus wide ‘storage architecture’ to support research use cases (e.g. instruments, data sharing) is required in the 2-5 year time window. - Instrumentation on campus is healthy and expanding. Technology must scale with this in the form of computation and storage. - Working with LEARN to upgrade network capacity (in multiples of 10G, or upgrades to 100G) will be required in the 1-3 year time frame. - Network monitoring and visibility will help to establish external science use cases. - Data sharing via portal systems is not currently a critical need, but growing in scope. EPOC can assist Baylor with options.

99 GENERAL AND MISCELLANEOUS↗

Effects of Plasma Etching on Dopant Compensation between p- and n-Type Poly-Si Fingers in Passivated Interdigitated Back Contact Solar Cells

Efficiencies surpassing 26% have been achieved for single-junction Si solar cells based on the interdigitated back contact (IBC) architecture. However, the possibility of lateral shunting between the doped fingers has limited industrial adoption thus far. To avoid this possibility of shunt, complicated patterning techniques of the doped rear fingers have been developed, but spreading can still occur through multiple pathways during cell processing. Patterning can be simplified by using masked plasma-enhanced chemical vapor deposition (PECVD), but spreading during deposition will also lead to contamination of the isolation region between doped fingers. We significantly reduce the effect of this spreading through a short, gentle plasma etching step, which enables a trap-assisted compensation mechanism to take effect more easily than without the plasma etch. These two effects combined allow for the use of a simple patterning technique while still maintaining a highly resistive region between the doped fingers, and will reduce the complexity of IBC cell fabrication overall.

atom probe↗

Two-dimensional shape recognition using sparse distributed memory

Researchers propose a method for recognizing two-dimensional shapes (hand-drawn characters, for example) with an associative memory. The method consists of two stages: first, the image is preprocessed to extract tangents to the contour of the shape; second, the set of tangents is converted to a long bit string for recognition with sparse distributed memory (SDM). SDM provides a simple, massively parallel architecture for an associative memory. Long bit vectors (256 to 1000 bits, for example) serve as both data and addresses to the memory, and patterns are grouped or classified according to similarity in Hamming distance. At the moment, tangents are extracted in a simple manner by progressively blurring the image and then using a Canny-type edge detector (Canny, 1986) to find edges at each stage of blurring. This results in a grid of tangents. While the technique used for obtaining the tangents is at present rather ad hoc, researchers plan to adopt an existing framework for extracting edge orientation information over a variety of resolutions, such as suggested by Watson (1987, 1983), Marr and Hildreth (1980), or Canny (1986).

Kanerva, Pentti↗

Detectors and Focal Plane Modules for Weather Instruments

Weather satellite instruments require detectors with a variety of wavelengths ranging from the visible to VLWIR. The Cross-track infrared Sounder (CrIS) is a Polar Orbiting interferometric sensor that measures earth radiances at high spectral resolution, using the data to provide pressure, temperature and moisture profiles of the atmosphere. The pressure, temperature and moisture sounding data are used in weather prediction models that track storms, predict levels of precipitation etc. The CrIS instrument contains SWIR (lambda(sub c) (is) approximately 5 micrometers at 98 K), MWIR (lambda(sub c) (is) approximately 9 micrometers at 98 K) and LWIRs (lambda(sub c) (is) approximately 15.4 m at 81 K) bands in three Focal Plane Array Assemblies (FPAAs). CrIS detectors are 850 micrometers diameter detectors with each FPAA consisting of nine photovoltaic detectors arranged in a 3 x 3 pattern. Molecular beam epitaxy (MBE)-grown Hg1-xCdxTe material are used for the detectors fabricated in a modified Double Layer Planar Heterostructure (DLPH) architecture. Each detector has an accompanying cold preamplifier. SWIR and MWIR FPAAs operate at 98 K and the LWIR FPAA at 81 K, permitting the use of passive radiators to cool the detectors. D* requirements at peak 14.01 micrometers wavelength are greater than 5.0E+10 Jones for LWIR, greater than 7.5E+10 Jones at 8.26 micrometers for MWIR and greater than 3.0E+11 Jones at peak 4.64 micrometers wavelength for SWIR. All FPAAs exceeded the D* requirements. Measured mean values for the nine photodiodes in each of the LWIR, MWIR and SWIR FPAAs are D* = 5.3 x 10(exp 10) cm-Hz1/2/W at 14.0 micrometers, 9.6 x 10(exp 10) cm-Hz1/2/W at 8.0 micrometers and 3.4 x 10(exp 11) cm-Hz1/2/W at 4.64 micrometers.

satellite instruments↗

Adapting In Situ Accelerators for Sparsity With Granular Matrix Reordering

Neural network (NN) inference is an essential part of modern systems and is found at the heart of numerous applications ranging from image recognition to natural language processing. In situ NN accelerators can efficiently perform NN inference using resistive crossbars, which makes them a promising solution to the data movement challenges faced by conventional architectures. Although such accelerators demonstrate significant potential for dense NNs, they often do not benefit from sparse NNs, which contain relatively few non-zero weights. Processing sparse NNs on in situ accelerators results in wasted energy to charge the entire crossbar where most elements are zeros. To address this limitation, this paper proposes Granular Matrix Reordering (GMR): a preprocessing technique that enables an energy-efficient computation of sparse NNs on in situ accelerators. GMR reorders the rows and columns of sparse weight matrices to maximize the crossbars' utilization and minimize the total number of crossbars needed to be charged. The reordering process does not rely on sparsity patterns and incurs no accuracy loss. Finally, GMR achieves an average of 28% and up to 34% reduction in energy consumption over seven pruned NNs across four different pruning methods and network architectures.

97 MATHEMATICS AND COMPUTING↗

Genetic architecture of leaf morphological and physiological traits in a Populus deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ pedigree revealed by quantitative trait locus analysis

Understanding the genetic architecture of leaf morphological and physiological traits will help plant breeders develop high biomass poplar genotypes. Quantitative trait locus (QTL) studies combining next-generation sequencing techniques can advance our understanding of the genetic basis of complex traits. In this study, we measured 13 leaf morphological and physiological traits and identified quantitative trait loci (QTLs) in a Populus deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ F1 population (500 progenies) using a high-density genetic map constructed by whole genome re-sequencing. This linkage map consisted of 5796 single nucleotide polymorphism (SNP) markers assigned to 19 linkage groups (LGs), spanning 2683.80 centimorgans (cM) of genetic length, with an average marker density of 0.46 cM. We identified 109 QTLs on 18 LGs for leaf morphological traits and 55 QTLs on 14 LGs for leaf physiological traits. One-hundred eight putative candidate genes were identified within the candidate genomic region. Co-expression network and gene ontology enrichment analyses suggested that these candidate genes were involved in the photosynthetic process. The differential expression patterns of the CYCLIN (Potri.015G112200) and RED CHLOROPHYLL REDUCTASE (Potri.007G043600) genes between two parents indicated their potential roles in leaf morphological and physiological traits. These findings decipher the genetic architecture of leaf morphological and physiological traits in the P. deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ pedigree and provide candidate genes for future poplar genetic improvement.

59 BASIC BIOLOGICAL SCIENCES↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Sensing Electrical Networks Securely & Economically (SENSE)

The growing adoption of distributed energy resources (DERs) like battery energy storage systems and roof top solar/PV and the rapid penetration of electric vehicles (EVs), the electric grid is undergoing a major transformation with elevated stress on legacy grid assets. Despite a lot of expenditure to address these challenges, both in dollars and manpower, utilities have not been able to receive the value that was promised. The gains have been most visible at the transmission and substation level, especially where the main objective was improving operational and economic efficiency for the utility. Improving visibility and control at a few select points enhances the existing and established paradigm of centralized command and control. With changing load patterns, load types and the overall transition to an “active grid”, the centralized control and coordination paradigm gets challenged. To address the challenges, a new architecture and mechanism is needed, one that supports decentralized control and decision making, extracting value streams at the grid edge, particularly as the changes are fueled by transitions occurring in the distribution system. To address this, a communications and data processing platform, “GAMMA” was developed and demonstrated through the project. At the heart of the platform, are distributed, intelligent edge nodes with sensing and compute capabilities, that can record and analyze information locally. They are embedded in sensors and actuators specific to different distribution system applications. Phase 1 of the project focused on developing novel sensor technology that can be used for monitoring utility pole top distribution transformers. The sensors were designed with the objective of being low-cost, communicating with the GAMMA cloud using novel “delay-tolerant” networking using Bluetooth and a secure mobile application. They were non-intrusive in nature so that they can be installed quickly in the field, resulting in overall low cost of deployment and operations. Following the successful completion of Phase 1, the team manufactured 100 units for a field demonstration in Phase 2. The field demonstration was carried out on two real feeder systems with the local utility partner. In total, 100 sensors were installed and operated over a period of 6 months in the state of Georgia. The platform is operational end to end, with the cloud infrastructure deployed on a distributed, serverless environment that can serve multiple data streams, an analytics engine and a portal to securely view the data from multiple assets. The data collected through the GAMMA Mobile Phone app showcased the viability of the novel delay tolerant networking architecture, and the data processing algorithms developed through the course of the project, were successful in extracting important information about the overall network, improving the utility’s visibility and situational awareness in the distribution feeder.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Kinetic Inductance Phonon-Mediated Detectors for Dark Matter: R&D at NEXUS

Superconducting thin films have long been used as phonon sensors, particularly in the field of dark matter (DM) direct detection, due to the meV-scale Cooper-pair binding energy. A novel class of these detectors based on microwave kinetic inductance detectors, dubbed Kinetic Inductance Phonon-Mediated Detectors (KIPMDs), offers an attractive architecture for microcalorimeters to probe DM down to the fermionic thermal relic mass limit of a few keV. Such a device featuring an aluminum resonator patterned onto a silicon substrate was operated at the NEXUS low-background facility at Fermilab for characterization and evaluation of its efficacy for a dark matter search. With this device we have demonstrated a resolution on the energy absorbed by the superconductor of 2.5 eV, a factor of two better than current state-of-the-art. In this talk, I will present our measurement of the energy resolution and phonon collection efficiency performed by exposing the bare substrate to a pulsed source of 470 nm photons. I will also discuss the path forward to obtaining in these devices the sub-eV resolution required to test the “Freeze-in” class of DM models. Finally, I will review other complementary efforts in our group to develop a superconducting qubit-based low-threshold detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Designing Energy-Efficient Quantum Computers Through Prediction and Reduction of Cooling Requirements for Cryogenic Electronics

Quantum computing has been identified as a “wild card” by the International Energy Agency in predicting future global data center energy usage. This is primarily because both uncertainty in the extent to which quantum computing will be adopted, and uncertainty in the power consumption of individual quantum data centers. Unlike the classical counterparts, quantum computers need to be maintained at near absolute zero, requiring energy-intensive cryogenic cooling systems. Therefore, as quantum computers scale up from existing 50 qubit technology demonstrations to the 10,000 to 100,000 qubit systems that will be able to solve complex problems, the energy consumption of both the electronics and the required cooling systems will also increase. To predict this scaling, this work analyzes the energy requirements for both computation and cooling of quantum hardware. We show that the energy requirements for cooling of quantum computers is determined by several computing system parameters, including the number and type of physical qubits, the operating temperature, the packaging efficiency of the system, and the split between circuits operating at cryogenic temperatures and those operating at room temperature. The energy requirements can then be found based on thermal system parameters such as cooling efficiency and cryostat heat transfer. Analysis of these parameters shows that the energy required for cooling is significantly larger than that required for computation, a reversal from energy usage patterns seen in conventional computing. The results and discussions provide a road-map for creating energy efficient quantum computers through the selection of computer architectures and cryogenic system configurations that minimize cooling requirements.

energy efficiency↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗