Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dynamic clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Model and remote-sensing-guided experimental design and hypothesis generation for monitoring snow-soil–plant interactions

In this study, we develop a machine-learning (ML)-enabled strategy for selecting hillslope-scale ecohydrological monitoring sites within snow-dominated mountainous watersheds, with a particular focus on snow-soil–plant interactions. Data layers rely on spatial data layers from both remote sensing and hydrological model simulations. Specifically, a Landsat-based foresummer drought sensitivity index is used to define the dependency of the annual peak plant productivity on the Palmer drought severity index in the early growing season. Hydrological simulations provide the spatiotemporal dynamics of near-surface soil moisture and snow depth. In this framework, a regression analysis identifies the key hydrological variables relevant to the spatial heterogeneity of drought sensitivity. We then apply unsupervised clustering to these key variables, using the Gaussian mixture model, to group hillslopes into several zones that have divergent relationships regarding soil moisture, snow dynamics, and drought sensitivity. Using the datasets collected in the East River Watershed (Crested Butte, Colorado, United States), results show that drought sensitivity is significantly correlated with model-derived soil moisture and snow-free timing over space and time. The relationship is, however, non-linear, such that the correlation decreases above a threshold elevation and in a heavy snow year due to large snowpacks, lateral flow, and soil storage limitations. Clustering is then able to define the zones that have high or low sensitivity to drought, as well as the mid-elevation regions where sensitivity is associated with the topographic aspect and net potential radiation. In addition, the algorithm identifies the most representative hillslopes with road/trail access within each zone for installing monitoring sites. Our method also aims to significantly increase the use of ML and model-simulation results to guide critical zone and watershed monitoring activities.

54 ENVIRONMENTAL SCIENCES↗

Time Dependence of Collision Probabilities During Satellite Conjunctions

The NASA Conjunction Assessment Risk Analysis (CARA) team has recently implemented updated software to calculate the probability of collision (P (sub c)) for Earth-orbiting satellites. The algorithm can employ complex dynamical models for orbital motion, and account for the effects of non-linear trajectories as well as both position and velocity uncertainties. This “3D P (sub c)” method entails computing a 3-dimensional numerical integral for each estimated probability. Our analysis indicates that the 3D method provides several new insights over the traditional “2D P (sub c)” method, even when approximating the orbital motion using the relatively simple Keplerian two-body dynamical model. First, the formulation provides the means to estimate variations in the time derivative of the collision probability, or the probability rate, R (sub c). For close-proximity satellites, such as those orbiting in formations or clusters, R (sub c) variations can show multiple peaks that repeat or blend with one another, providing insight into the ongoing temporal distribution of risk. For single, isolated conjunctions, R (sub c) analysis provides the means to identify and bound the times of peak collision risk. Additionally, analysis of multiple actual archived conjunctions demonstrates that the commonly used “2D P (sub c)” approximation can occasionally provide inaccurate estimates. These include cases in which the 2D method yields negligibly small probabilities (e.g., P (sub c)) is greater than 10 (sup -10)), but the 3D estimates are sufficiently large to prompt increased monitoring or collision mitigation (e.g., P (sub c) is greater than or equal to 10 (sup -5)). Finally, the archive analysis indicates that a relatively efficient calculation can be used to identify which conjunctions will have negligibly small probabilities. This small-P (sub c) screening test can significantly speed the overall risk analysis computation for large numbers of conjunctions.

Hall, Doyle T.↗

Toward GEOS-6, A Global Cloud System Resolving Atmospheric Model

NASA is committed to observing and understanding the weather and climate of our home planet through the use of multi-scale modeling systems and space-based observations. Global climate models have evolved to take advantage of the influx of multi- and many-core computing technologies and the availability of large clusters of multi-core microprocessors. GEOS-6 is a next-generation cloud system resolving atmospheric model that will place NASA at the forefront of scientific exploration of our atmosphere and climate. Model simulations with GEOS-6 will produce a realistic representation of our atmosphere on the scale of typical satellite observations, bringing a visual comprehension of model results to a new level among the climate enthusiasts. In preparation for GEOS-6, the agency's flagship Earth System Modeling Framework [JDl] has been enhanced to support cutting-edge high-resolution global climate and weather simulations. Improvements include a cubed-sphere grid that exposes parallelism; a non-hydrostatic finite volume dynamical core, and algorithm designed for co-processor technologies, among others. GEOS-6 represents a fundamental advancement in the capability of global Earth system models. The ability to directly compare global simulations at the resolution of spaceborne satellite images will lead to algorithm improvements and better utilization of space-based observations within the GOES data assimilation system

Putman, William M.↗

Extracting Independent Local Oscillatory Geophysical Signals by Geodetic Tropospheric Delay

Zenith Tropospheric Delay (ZTD) due to water vapor derived from space geodetic techniques and numerical weather prediction simulated-reanalysis data exhibits non-linear and non-stationary properties akin to those in the crucial geophysical signals of interest to the research community. These time series, once decomposed into additive (and stochastic) components, have information about the long term global change (the trend) and other interpretable (quasi-) periodic components such as seasonal cycles and noise. Such stochastic component(s) could be a function that exhibits at most one extremum within a data span or a monotonic function within a certain temporal span. In this contribution, we examine the use of the combined Ensemble Empirical Mode Decomposition (EEMD) and Independent Component Analysis (ICA): the EEMD-ICA algorithm to extract the independent local oscillatory stochastic components in the tropospheric delay derived from the European Centre for Medium-Range Weather Forecasts (ECMWF) over six geodetic sites (HartRAO, Hobart26, Wettzell, Gilcreek, Westford, and Tsukub32). The proposed methodology allows independent geophysical processes to be extracted and assessed. Analysis of the quality index of the Independent Components (ICs) derived for each cluster of local oscillatory components (also called the Intrinsic Mode Functions (IMFs)) for all the geodetic stations considered in the study demonstrate that they are strongly site dependent. Such strong dependency seems to suggest that the localized geophysical signals embedded in the ZTD over the geodetic sites are not correlated. Further, from the viewpoint of non-linear dynamical systems, four geophysical signals the Quasi-Biennial Oscillation (QBO) index derived from the NCEP/NCAR reanalysis, the Southern Oscillation Index (SOI) anomaly from NCEP, the SIDC monthly Sun Spot Number (SSN), and the Length of Day (LoD) are linked to the extracted signal components from ZTD. Results from the synchronization analysis show that ZTD and the geophysical signals exhibit (albeit subtle) site dependent phase synchronization index.

Botai, O. J.↗

Machine-learning identification of the variability of mean velocity and turbulence intensity for wakes generated by onshore wind turbines: Cluster analysis of wind LiDAR measurements

Light detection and ranging (LiDAR) measurements of isolated wakes generated by wind turbines installed at an onshore wind farm are leveraged to characterize the variability of the wake mean velocity and turbulence intensity during typical operations, which encompass a breadth of atmospheric stability regimes and rotor thrust coefficients. The LiDAR measurements are clustered through the k-means algorithm, which enables identifying the most representative realizations of wind turbine wakes while avoiding the imposition of thresholds for the various wind and turbine parameters. Considering the large number of LiDAR samples collected to probe the wake velocity field, the dimensionality of the experimental dataset is reduced by projecting the LiDAR data on an intelligently truncated basis obtained with the proper orthogonal decomposition (POD). The coefficients of only five physics-informed POD modes are then injected in the k-means algorithm for clustering the LiDAR dataset. The analysis of the clustered LiDAR data and the associated supervisory control and data acquisition and meteorological data enables the study of the variability of the wake velocity deficit, wake extent, and wake-added turbulence intensity for different thrust coefficients of the turbine rotor and regimes of atmospheric stability. Furthermore, the cluster analysis of the LiDAR data allows for the identification of systematic off-design operations with a certain yaw misalignment of the turbine rotor with the mean wind direction.

17 WIND ENERGY↗

SV-Sim: Scalable PGAS-based State Vector Simulation of Quantum Circuits

High-performance quantum circuit simulation in a classic HPC is still imperative in the NISQ era. Observing that the major obstacle of scalable state-vector quantum simulation arises from the massively fine-grained irregular data-exchange with remote nodes, in this paper we present SV-Sim to apply the emerging PGAS-based communication models (i.e., direct peer access for intra-node CPUs/GPUs and SHMEM for inter-node CPU/GPU clusters) for efficient scalable quantum circuit simulation. Through an orchestrated device functional pointer design, SV-Sim is able to abstract the quantum gate sets across various heterogeneous backends, including IBM/Intel/AMD CPUs, NVIDIA /AMD GPUs, and Intel MIC, in a unified framework, but still asserting outstanding performance and tractable interface to higher-level quantum programming environments, such as IBM Qiskit, Microsoft Q\# and Google Cirq. Circumventing the disability of polymorphism in GPUs and leveraging the device-initiated one-sided communication, SV-Sim can process dynamically synthesized quantum circuit in a single GPU/CPU kernel without the need of expensive JIT or runtime branching, significantly improving the performance and simplifying the programming complexity for the emerging variational quantum algorithms. Evaluations on NVIDIA A100-DGX-1, V100-DGX-2, AMD MI100, ALCF Theta, and OLCF Summit HPCs show that SV-Sim can delivery scalable performance on various state-of-the-art HPC platforms, offering a useful tool for quantum algorithm validation and verification.

Li, Ang↗

Revisiting single inclusive jet production: timelike factorization and reciprocity

Factorization theorems for single inclusive jet production play a crucial role in the study of jets and their substructure. In the case of small radius jets, the dynamics of the jet clustering can be factorized from both the hard production dynamics, and the dynamics of the low scale jet substructure measurement, and is described by a matching coefficient that can be computed in perturbative Quantum Chromodynamics (QCD). A proposed factorization formula describing this process has been previously presented in the literature, and is referred to as the semi-inclusive, or fragmenting jets formalism. By performing an explicit two-loop calculation, we show the inconsistency of this factorization formula, in agreement with another recent result in the literature. Building on recent progress in the factorization of single logarithmic observables, and the understanding of reciprocity, we then derive a new all-order factorization theorem for inclusive jet production. The use of a jet algorithm, being only a modification of the infrared structure of the measurement, modifies the structure of convolutions in the factorization theorem, as compared to inclusive fragmentation, but maintains the universality of the inclusive hard function and its associated Dokshitzer-Gribov-Lipatov-Altarelli-Parisi (DGLAP) evolution, which are ultraviolet properties. However, the non-trivial structure of convolutions in the factorization theorem implies that the jet functions exhibit a modified evolution. We perform an explicit two-loop calculation of the jet function in both N = 4 super Yang-Mills (SYM), and for all color channels in QCD, finding exact agreement with the structure derived from our renormalization group equations. In addition, we derive several new results, including an extension of our factorization formula to jet substructure observables, a jet algorithm definition of a generating function for the energy correlators, and new results for exclusive jet functions. Our results are a key ingredient for achieving precision jet substructure at colliders.

Effective Field Theories↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 kin or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed- shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

NASA Tech Briefs, August 2008

Customizable Digital Receivers for Radar Two-Camera Acquisition and Tracking of a Flying Target Visual Data Analysis for Satellites A Data Type for Efficient Representation of Other Data Types Hand-Held Ultrasonic Instrument for Reading Matrix Symbols Broadband Microstrip-to-Coplanar Strip Double-Y Balun A Topographical Lidar System for Terrain-Relative Navigation Programmable Low-Voltage Circuit Breaker and Tester Electronic Switch Arrays for Managing Microbattery Arrays Topics covered include: Lower-Dark-Current, Higher-Blue-Response CMOS Imagers; Fabricating Large-Area Sheets of Single-Layer Graphene by CVD; Support for Diagnosis of Custom Computer Hardware; Providing Goal-Based Autonomy for Commanding a Spacecraft; Dynamic Method for Identifying Collected Sample Mass; Optimal Planning and Problem-Solving; Attitude-Control Algorithm for Minimizing Maneuver Execution Errors; Grants Document-Generation System; Heat-Storage Modules Containing LiNO3 3H2O and Graphite Foam; Precipitation-Strengthened, High-Temperature, High-Force Shape Memory Alloys; Improved Relief Valve Would Be Less Susceptible to Failure; Safety Modification of Cam-and-Groove Hose Coupling; Using Composite Materials in a Cryogenic Pump; Using Electronic Noses to Detect Tumors During Neurosurgery; Producing Newborn Synchronous Mammalian Cells; Smaller, Lower-Power Fast-Neutron Scintillation Detectors; Rotationally Vibrating Electric-Field Mill; Estimating Hardness from the USDC Tool-Bit Temperature Rise; Particle-Charge Spectrometer; Automated Production of Movies on a Cluster of Computers; FIDO-Class Development Rover; and Tone-Based Command of Deep Space Probes Using Ground Antennas.

Source record↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 km or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed-shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

Modeling of H2 Dispersion at ARIES

Hydrogen is a versatile and clean energy carrier that can be produced from various renewable sources such as wind, solar, and hydropower and help decarbonize electricity grids, industry, and transportation. Using the Hydrogen Research Facility under Advanced Research on Integrated Energy Systems (ARIES) at the National Renewable Energy Laboratory's (NREL) Flatirons campus as a test bench, the study examines the feasibility, useability, and value of using computational fluid dynamics (CFD) techniques to model hydrogen dispersion. The ARIES facility was chosen because controlled hydrogen releases can be performed at a rate of 27 kg-H2/hr. Site-specific atmospheric and weather condition data such as wind speed and temperature were used as inputs to the model. The results show statistical distributions and ranges of hydrogen concentrations at locations throughout the domain. Wind conditions are found to significantly impact the release behavior, including the hydrogen cloud's direction and concentrations. At low wind speeds (below 1 mph), hydrogen forms a cloud and at higher wind speeds (> 2-4 mph) hydrogen plume stretches in the direction of wind momentum. From >100 simulations for ARIES site-specific conditions, statistical quantities combined with a clustering algorithm were used to propose sensor location at various elevations from ground.

dispersion↗

A Global End-Member Approach to Derive aCDOM(440) from Near-Surface Optical Measurements

This study establishes an optical inversion scheme for deriving the absorption coefficient of colored (or chromophoric, depending on the literature) dissolved organic material (CDOM) at the 440 nm wavelength, which can be applied to global water masses with near-equal efficacy. The approach uses a ratio of diffuse attenuation coefficient spectral end members, i.e., a short and long wavelength pair. The global perspective is established by sampling "extremely" clear water plus a generalized extent in turbidity and optical properties that each span three decades of dynamic range. A unique data set was collected in oceanic, coastal, and inland waters (as shallow as 0.6 m) from the North Pacific Ocean, the Arctic Ocean, Hawaii, Japan, Puerto Rico, and the east and west coasts of the United States. The data were partitioned using subjective categorizations to define a validation quality subset of conservative water masses, i.e., the inflow and outflow of properties constrain the range in the gradient of a constituent, plus 15 subcategories of water masses that were not evolving conservatively. The dependence on subcategories was confirmed with an objective methodology based on cluster analysis techniques. The latter defined five distinct classes with validation quality data present in all classes, but which also decreased in percent composition as a function of increasing class number and optical complexity. Four different algorithms based on different validation quality end members were validated with accuracies of 1.–6.2 %, wherein pairs with the largest spectral span were most accurate. Although algorithm accuracy decreased with the inclusion of more subcategories containing non-conservative water masses, changes to the algorithm fit were small when a preponderance of subcategories were included. The high accuracy for all end-member algorithms was the result of data acquisition and data processing improvements, e.g., increased vertical sampling resolution to less than 1mm and a boundary constraint to mitigate wave focusing effects, respectively. An independent evaluation with a historical database confirmed the consistency of the algorithmic approach and its application to quality assurance, e.g., to flag data outside expected ranges, identify suspect spectra, and objectively determine the in-water extrapolation interval by converging agreement for all applicable end-member algorithms. The legacy data exhibit degraded performance (as 44 % uncertainty) due to a lack of high-quality near-surface observations, especially for clear waters wherein wave-focusing effects are problematic. The novel optical approach allows the in situ estimation of an in-water constituent in keeping with the accuracy obtained in the laboratory.

Stanford B Hooker↗

The void spectrum in two-dimensional numerical simulations of gravitational clustering

An algorithm for deriving a spectrum of void sizes from two-dimensional high-resolution numerical simulations of gravitational clustering is tested, and it is verified that it produces the correct results where those results can be anticipated. The method is used to study the growth of voids as clustering proceeds. It is found that the most stable indicator of the characteristic void 'size' in the simulations is the mean fractional area covered by voids of diameter d, in a density field smoothed at its correlation length. Very accurate scaling behavior is found in power-law numerical models as they evolve. Eventually, this scaling breaks down as the nonlinearity reaches larger scales. It is shown that this breakdown is a manifestation of the undesirable effect of boundary conditions on simulations, even with the very large dynamic range possible here. A simple criterion is suggested for deciding when simulations with modest large-scale power may systematically underestimate the frequency of larger voids.

Kauffmann, Guinevere↗

Scheduling Operations for Massive Heterogeneous Clusters

High-performance computing (HPC) programming has become increasingly difficult with the advent of hybrid supercomputers consisting of multicore CPUs and accelerator boards such as the GPU. Manual tuning of software to achieve high performance on this type of machine has been performed by programmers. This is needlessly difficult and prone to being invalidated by new hardware, new software, or changes in the underlying code. A system was developed for task-based representation of programs, which when coupled with a scheduler and runtime system, allows for many benefits, including higher performance and utilization of computational resources, easier programming and porting, and adaptations of code during runtime. The system consists of a method of representing computer algorithms as a series of data-dependent tasks. The series forms a graph, which can be scheduled for execution on many nodes of a supercomputer efficiently by a computer algorithm. The schedule is executed by a dispatch component, which is tailored to understand all of the hardware types that may be available within the system. The scheduler is informed by a cluster mapping tool, which generates a topology of available resources and their strengths and communication costs. Software is decoupled from its hardware, which aids in porting to future architectures. A computer algorithm schedules all operations, which for systems of high complexity (i.e., most NASA codes), cannot be performed optimally by a human. The system aids in reducing repetitive code, such as communication code, and aids in the reduction of redundant code across projects. It adds new features to code automatically, such as recovering from a lost node or the ability to modify the code while running. In this project, the innovators at the time of this reporting intend to develop two distinct technologies that build upon each other and both of which serve as building blocks for more efficient HPC usage. First is the scheduling and dynamic execution framework, and the second is scalable linear algebra libraries that are built directly on the former.

Humphrey, John↗

Short- and medium-range orders in Al90Tb10 glass and their relation to the structures of competing crystalline phases

Molecular dynamics simulations using an interatomic potential developed by artificial neural network deep machine learning are performed to study the local structural order in Al90Tb10 metallic glass. We show that more than 80% of the Tb-centered clusters in Al90Tb10 glass have short-range order (SRO) with their 17 first coordination shell atoms stacked in a ‘3661’ or ‘15551’ sequence. Medium-range order (MRO) in Bergman-type packing extended out to the second and third coordination shells is also clearly observed. Analysis of the network formed by the ‘3661’ and ‘15551’ clusters show that ~82% of such SRO units share their faces or vertexes, while only ~6% of neighboring SRO pairs are interpenetrating. Such a network topology is consistent with the Bergman-type MRO around the Tb-centers. Moreover, crystal structure searches using genetic algorithm and the neural network interatomic potential reveal several low-energy metastable crystalline structures in the composition range close to Al 90 Tb 10 . Some of these crystalline structures have the ‘3661’ SRO while others have the ‘15551’ SRO. While the crystalline structures with the ‘3661’ SRO also exhibit the MRO very similar to that observed in the glass, the ones with the ‘15551’ SRO have very different atomic packing in the second and third shells around the Tb centers from that of the Bergman-type MRO observed in the glassy phase.

36 MATERIALS SCIENCE↗

Continuous-variable quantum computation of the O(3) model in 1+1 dimensions

We formulate the $O(3)$ non-linear sigma model in $1+1$ dimensions as a limit of a three-component scalar field theory restricted to the unit sphere in the large squeezing limit. This allows us to describe the model in terms of the continuous variable (CV) approach to quantum computing. Here we construct the ground state and excited states using the coupled cluster ansatz and find excellent agreement with the exact diagonalization results for a small number of lattice sites. We then present the simulation protocol for the time evolution of the model using CV gates, estimate the discretization error, and present numerical results obtained from a photonic quantum simulator. We expect that the methods developed in this work will be useful for exploring interesting dynamics for a wide class of sigma models and gauge theories, as well as for simulating scattering events on quantum hardware in the coming decade.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗