Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Algorithms and data structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Boosting Barlow Twins Reduced Order Modeling for Machine Learning‐Based Surrogate Models in Multiphase Flow Problems

Abstract We present an innovative approach called boosting Barlow Twins reduced order modeling (BBT‐ROM) to enhance the reliability of machine learning surrogate models for multiphase flow problems. BBT‐ROM builds upon Barlow Twins reduced order modeling that leverages self‐supervised learning to effectively handle linear and nonlinear manifolds by constructing well‐structured latent spaces of input parameters and output quantities. To address the challenge of high contrast data in multiphase flow problems due to injection wells and faults, we employ a boosting algorithm within BBT‐ROM. This algorithm sequentially trains a set of weak models (i.e., inaccurate models), improving prediction accuracy through ensemble learning. To evaluate the performance of BBT‐ROM, we conduct three three‐dimensional multiphase flow problems, including waterflooding and geologic carbon storage (GCS), with varying numbers of input parameter cases and model domain features. The results demonstrate that BBT‐ROM excels at predicting non‐wetting phase saturation (e.g., oil or saturation) and fluid pressure, with average relative errors ranging from 0.5% to 3%. Importantly, BBT‐ROM showcases robustness when faced with limited input parameter space during GCS testing.

58 GEOSCIENCES↗

HIV drug resistance during antiretroviral therapy scale-up in Uganda, 2012–19: a population-based, longitudinal study

Background With scale-up of antiretroviral therapy (ART) in sub-Saharan Africa, increasing pretreatment HIV drug resistance has been reported; however, the broader effect of ART expansion on population-level resistance patterns remains insufficiently quantified. We aimed to estimate the longitudinal prevalence of drug resistance and resistance-conferring mutations. Methods This study used data collected as part of the Rakai Community Cohort Study (RCCS), an open population-based census and cohort study conducted in southern Uganda. At each survey round, residents aged 15–49 years are invited to participate and receive a structured questionnaire that obtains sociodemographic, behavioural, and health information, including self-reported past and current ART use. Voluntary HIV testing is conducted using a rapid test algorithm and a venous blood sample. People with HIV provide samples for viral load quantification and deep sequencing. We analysed RCCS survey, HIV viral load, and deep sequencing (which was used to predict resistance) data from five survey rounds. The key outcomes were the population prevalence of viraemic people with HIV with non-nucleoside reverse transcriptase inhibitor (NNRTI), nucleoside reverse transcriptase inhibitor (NRTI), protease inhibitor, or multiclass resistance among all participants (regardless of HIV serostatus) in the 2015 and 2017 surveys. Prevalence of class-specific resistance and resistance-conferring substitutions were estimated using robust log-Poisson regression. Findings Between Aug 10, 2011, and Nov 4, 2020, there were 43 361 participants in the RCCS and 7923 (18·27%) people with HIV. Over five survey rounds, 93 622 participant visits occurred, among which 17 460 (18·65%) were from people with HIV. Over the analysis period, the median age of study participants remained similar (28 years [22–35] in 2012 and 29 years [21–38] in 2019). Sufficient data were available to reliably genotype 4072 (90·03%) of 4523 participant visits from 3407 people with HIV for at least one drug. Overall population prevalence of resistance contributed by viraemic pretreatment people with HIV decreased between 2012 and 2017 from 0·56% (95% CI 0·42–0·75) to 0·25% (0·18–0·33) for NNRTI and from 0·24% (0·15–0·37) to 0·05% (0·02–0·10) for NRTI (prevalence ratio 0·44 [0·29–0·68] for NNRTI and 0·21 [0·09–0·47] for NRTI). Between 2012 and 2017, NNRTI resistance among viraemic pretreatment people with HIV increased from 4·86% (3·69–6·42) to 9·61% (7·27–12·7; prevalence ratio 1·98 [1·34–2·91]). The prevalence of NNRTI and NRTI resistance was substantially higher among viraemic treatment-experienced people with HIV (51·49% [46·24–57·34] for NNRTI and 36·46% [30·06–44·22] for NRTI in 2017) than among pretreatment people with HIV. NNRTI and NRTI resistance was predominantly attributable to rtK103N and rtM184V. inT97A was observed at a similar prevalence among viraemic treatment-experienced (9·96% [6·41–15·48]) and viraemic pretreatment (10·56% [8·01–13·93]) people with HIV; no major dolutegravir resistance mutations were observed. Interpretation Despite rising NNRTI resistance among pretreatment people with HIV, overall population prevalence of pretreatment HIV drug-resistant viraemia decreased due to increasing ART uptake and viral suppression. This finding underscores the crucial role of achieving and maintaining high ART coverage in reducing transmission of drug-resistant HIV. The high prevalence of mutations conferring resistance to components of first-line ART regimens among viraemic people with HIV is potentially concerning. Funding National Institutes of Health, Johns Hopkins University Center for AIDS Research, Bill & Melinda Gates Foundation, and the US Centers for Disease Control and Prevention.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Augmentation, Hypersweeps, and Branch Decomposition of Contour Trees for Scientific Exploration

Contour trees describe the topology of level sets in scalar fields and are widely used in topological data analysis and visualization. A main challenge of utilizing contour trees for large-scale scientific data is their computation at scale using highperformance computing. To address this challenge, recent work has introduced distributed hierarchical contour trees for distributed computation and storage of contour trees. However, effective use of these distributed structures in analysis and visualization requires subsequent computation of geometric properties and branch decomposition to support contour extraction and exploration. In this work, we introduce distributed algorithms for augmentation, hypersweeps, and branch decomposition that enable parallel computation of geometric properties, and support the use of distributed contour trees as query structures for scientific exploration. Finally, we evaluate the parallel performance of these algorithms and apply them to identify and extract important contours for scientific visualization.

97 MATHEMATICS AND COMPUTING↗

Cross Section Evaluation for Exclusive Channels of K+Λ and K+Σ0 Electroproduction off Protons Using CLAS Detector Data

In this work, a method for evaluating the cross sections of electroproduction of K+Λ0 and K+Σ0 off protons in the region of invariant masses of final hadrons MK + MY <W < 2.65 GeV (MK and MY being the masses of the kaon and hyperon, respectively) and squares of four-momentum transfers of virtual photons, i.e., photon virtualities 0 < Q2 < 5 GeV2, is developed based on experimental data of these exclusive channels’ cross sections measured by the CLAS detector in Hall B at Jefferson Lab. A set of algorithms has been implemented to evaluate the differential cross sections of these channels, along with their statistical and systematic uncertainties. A program was developed for the evaluation of differential cross sections and structure functions using C++ and Python libraries. An interactive website was created for working with the program, enabling the analysis of one-dimensional and two-dimensional dependences of structure functions and differential cross sections. The evaluation of differential cross sections for the K+Λ and K+Σ0 electroproduction channels is necessary for extracting the structure function σLT from the data on the polarization asymmetry of electroproduction reactions of these final states with longitudinally polarized electrons. The obtained results are also important for the development of realistic Monte Carlo event generators in planning future experiments and for evaluating the efficiency of detecting final particles when extracting reaction cross sections from experimental data.

Golda, A. V.↗

Revealing the Structure and Dynamics of Self-Generated Electric and Magnetic Fields Near Plasma Stagnation in Laser-Driven Hohlraums

By coupling newly developed triparticle charged particle radiography with radiography reconstruction algorithms and novel reconstruction postprocessing techniques, the spatial structure and time evolution of self-generated electric and magnetic fields in laser-driven vacuum hohlraums have been quantitatively revealed. Through high-fidelity data from a series of experiments, it is shown that late in the hohlraum evolution (after the end of laser drive) these fields are strongly correlated in both space and time, providing evidence that their evolution is primarily dominated by advection with the plasma flow. At these late times, plasma flow velocities inferred from both gross radiography analysis and field reconstructions (and corroborated with Thomson scattering measurements) indicate that the plasma is approaching stagnation near the hohlraum axis. Finally, these experiments provide not only new physical insight into spontaneously generated hohlraum fields, but also provide important spatially and temporally resolved information for future benchmarking of numerical codes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion↗

A Computational Framework to design 3D stiffness gradient acoustic metamaterials for impedance matching

Acoustic waves play a crucial role in various applications, including medical imaging, non-destructive testing, and sonar systems. One of the significant challenges in these applications is impedance matching, which is essential for minimizing reflections and maximizing the transfer of acoustic energy between different media. Acoustic metamaterials offer a promising solution to this challenge. In addition to impedance control, gradient stiffness can enhance structural efficiency and enable spatial control of wave propagation, making it a valuable feature in acoustic metamaterial design. In this pa- per, we present our developed computational method to design 3D stiffness gradient acoustic metamaterials for impedance matching. The key steps in our approach include generating initial designs using a periodic covariance function to provide unit cells that are both periodic on the boundaries and randomly formed inside the unit cell. Furthermore, we integrated manufacturing constraints into the design process, ensuring that the structures are interconnected for fabrication. We propose two computational optimization algorithms: GenUnit, based on a non-dominated sorting genetic algorithm (NSGA-II), and MLMatch, which leverages differentiable machine learning. The two approaches are not separate contributions but complementary com- ponents of a unified framework. GenUnit requires no training data and directly interfaces with physics-based simulations, making it highly accurate but slower for large-scale exploration. In contrast, MLMatch is data-hungry during training but, once trained, enables near-instantaneous inference and broad design-space coverage. Together, they form a hybrid strategy: ML- Match rapidly explores the global design space, and GenUnit provides local refinement with high-fidelity accuracy. This balance between training cost, inference time, and precision is the motivation for including both methods in the same study. We applied this dual-algorithm framework to generate two metallic-based metamaterial designs that match the acoustic impedance of water while exhibiting a controlled gradient in stiffness (from stiff to soft). The stiffness gradient is particularly advantageous in applications where one side of the structure must interface with soft or sensitive surfaces, such as human tissue or delicate components. Here, this work paves the way for improved materials in various acoustic applications, particularly in ultrasound devices, by providing better impedance.

Metamaterial↗

guppy i : a code for reducing the storage requirements of cosmological simulations

ABSTRACT As cosmological simulations have grown in size, the permanent storage requirements of their particle data have also grown. Even modest simulations present a major logistical challenge for the groups which run these boxes and researchers without access to high performance computing facilities often need to restrict their analysis to lower quality data. In this paper, we present guppy, a compression algorithm and code base tailored to reduce the sizes of dark matter-only cosmological simulations by approximately an order of magnitude. guppy is a ‘lossy’ algorithm, meaning that it injects a small amount of controlled and uncorrelated noise into particle properties. We perform extensive tests on the impact that this noise has on the internal structure of dark matter haloes, and identify conservative accuracy limits which ensure that compression has no practical impact on single-snapshot halo properties, profiles, and abundances. We also release functional prototype libraries in C, Python, and Go for reading and creating guppy data.

79 ASTRONOMY AND ASTROPHYSICS↗

Materials Learning Algorithms (MALA): Scalable machine learning for electronic structure calculations in large-scale atomistic simulations

We present the Materials Learning Algorithms (MALA) package, a scalable machine learning framework designed to accelerate density functional theory (DFT) calculations suitable for large-scale atomistic simulations. Using local descriptors of the atomic environment, MALA models efficiently predict key electronic observables, including local density of states, electronic density, density of states, and total energy. The package integrates data sampling, model training and scalable inference into a unified library, while ensuring compatibility with standard DFT and molecular dynamics codes. We demonstrate MALA's capabilities with examples including boron clusters, aluminum across its solid-liquid phase boundary, and predicting the electronic structure of a stacking fault in a large beryllium slab. Scaling analyses reveal MALA's computational efficiency and identify bottlenecks for future optimization. With its ability to model electronic structures at scales far beyond standard DFT, MALA is well suited for modeling complex material systems, making it a versatile tool for advanced materials research.

Density functional theory↗

Adaptive anomaly detection for identifying attacks in cyber-physical systems: A systematic literature review

Modern cyberattacks in cyber-physical systems (CPS) rapidly evolve and cannot be deterred effectively with most current methods, which focus on characterizing past threats. Adaptive anomaly detection (AAD) is among the most promising techniques to detect evolving cyberattacks, with an emphasis on fast data processing and model adaptation. AAD has been researched extensively; however, to the best of our knowledge, our work is the first systematic literature review (SLR) on current research in this field. We present a comprehensive SLR, gathering 397 relevant papers and systematically analyzing 65 of them (47 research and 18 survey papers) on AAD in CPS from 2013 to November 2023. We introduce a novel taxonomy considering attack types, CPS application, learning paradigm, data management, and algorithms. Our findings show that most studies addressed either model adaptation or data processing, but rarely both simultaneously. This indicates a research gap in fully adaptive solutions. We also categorize algorithms, datasets, and attack characteristics, and summarize strengths and weaknesses across the literature. Our review provides a structured and accessible reference for researchers and practitioners, offering insights into key trends and highlighting limitations in current approaches. Finally, we outline several future research directions, including the need for integrated real-time processing and adaptive learning, explainability, and uncertainty quantification in AAD for CPS.

Adaptation↗

Neural chaos: A spectral stochastic neural operator

Building surrogate models for operators with uncertainty quantification capabilities is essential for many engineering applications where randomness–such as variability in material properties, boundary conditions, and initial conditions–is unavoidable. Polynomial Chaos Expansion (PCE) is widely recognized as a go-to method for constructing stochastic surrogates in both intrusive and non-intrusive ways, and it has recently been used in the context of operator learning. However, its application becomes challenging for complex or high-dimensional processes, as achieving accuracy requires higher-order polynomials, which can increase computational demand and/or the risk of overfitting. Furthermore, PCE requires specialized treatments to manage random variables that are not independent, and these treatments may be problem-dependent or may fail with increasing complexity. Here, in this work, we adopt the same formalism as the spectral expansion used in PCE; however, we replace the classical polynomial basis functions with neural network (NN) basis functions to leverage their expressivity. To achieve this, we propose an algorithm that identifies NN-parameterized basis functions in a purely data-driven manner, without any prior assumptions about the joint distribution of the random variables involved, whether independent or dependent, or about their marginal distributions. The proposed algorithm identifies each NN-parameterized basis function sequentially, ensuring they are orthogonal with respect to the data distribution. The basis functions are constructed directly on the joint stochastic variables without requiring a tensor product structure or assuming independence of the random variables. This approach may offer greater flexibility for complex stochastic models, while simplifying implementation compared to the tensor product structures typically used in PCE to handle random vectors. This is particularly advantageous given the current state of open-source packages, where building and training neural networks can be done with just a few lines of code and extensive community support. We demonstrate the effectiveness of the proposed scheme through several numerical examples of varying complexity and provide comparisons with classical PCE.

Polynomial chaos expansion↗

Improving Property Graph Layouts by Leveraging Attribute Similarity for Structurally Equivalent Nodes

Many real-world networks contain structurally-equivalent nodes. These are defined as vertices that share the same set of neighboring nodes, making them interchangeable with a traditional graph layout approach. However, many real-world graphs also have properties associated with nodes, adding additional meaning to them. We present an approach for swapping locations of structurally-equivalent nodes in graph layout so that those with more similar properties have closer proximity to each other. This improves the usefulness of the visualization from an attribute perspective without negatively impacting the visualization from a structural perspective. We include an algorithm for finding these sets of nodes in linear time, as well as methodologies for ordering nodes based on their attribute similarity, which works for scalar, ordinal, multidimensional, and categorical data.

graph drawing, network visualization, property gra↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

A Bayesian desmearing algorithm for Bonse–Hart USANS with anisotropic scattering

Ultra-small-angle neutron scattering (USANS) using Bonse–Hart optics provides micrometer-scale structural insights but suffers from severe slit-geometry smearing. While well-established for isotropic systems, quantitative desmearing of anisotropic data remains a challenge because conventional corrections break down for non-radial scattering. In this work, we address this by developing a resolution-aware Bayesian framework that explicitly incorporates anisotropy via an affine deformation to the scattering pattern, guided by the principle of parsimony. This results in orientation-resolved point-spread functions that enable a self-consistent determination of both the resolution and deformation parameters. Using Gaussian process regression with uncertainty quantification and a probabilistic correction for multiple scattering, we demonstrate the framework’s effectiveness through numerical benchmarks and experimental studies of a stretched polymer melt. Our approach enables the seamless integration of SANS and USANS data, facilitating quantitative structural analysis of deformed materials at nanometer to micrometer scales.

36 MATERIALS SCIENCE↗

The DESI DR1 peculiar velocity survey: growth rate measurements from the maximum likelihood fields method

We present the constraint on the growth rate of structure from the combination of DESI DR1 BGS sample, Fundamental Plane, and Tully-Fisher peculiar velocity catalogues using the maximum likelihood fields method. The combined catalogue contains 415,523 galaxy redshifts and 76,616 peculiar velocity measurements. To handle the large amount of data in the DESI DR1 peculiar velocity catalogue, we significantly improve the computational efficiency by rewriting the algorithm with JAX. After removing outliers and Tully-Fisher galaxies that are affected by systematics, we find fσ 8 = 0.483 -0.043 +0.080 (stat) ± 0.018(sys), consistent within 1σ with the power spectrum and correlation function analysis using the same dataset. Combining all three measurements with appropriate correlations, the consensus measurement is fσ 8 (z eff = 0.07) = 0.450±0.055, consistent with Planck +ΛCDM cosmology (fσ 8 = 0.449±0.008). Combining with the high redshift growth rate of structure measurements from DESI ShapeFit, the constraint on the growth index is γ = 0.58±0.11, consistent with GR.

cosmic flows↗

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

FAIR Data and Interpretable AI Framework for Architectured Metamaterials

Our interdisciplinary effort successfully generated FAIR (Findable, Accessible, Interoperable, and Reusable) benchmark datasets for mechanical metamaterials while introducing a novel Artificial Intelligence (AI) framework known as Learning Refined Compositional Rules (LRCR). This framework was specifically designed to bridge the gap across varying computational length scales and extract the underlying physical mechanisms that connect a material's structural geometry to its bulk acoustic properties. Historically, the discovery of such structured materials relied heavily on human intuition or opaque, black-box optimization algorithms that were difficult to generalize. By combining interpretable machine learning techniques with rigorous experimental validation, this project established clear, generalizable design guidelines for tuning wave dispersion and controlling vibrations. Ultimately, the public availability of these structured datasets and algorithms will significantly reduce computational costs and accelerate the design of advanced multi-functional acoustic devices, offering broad societal impacts across fields like aerospace engineering, telecommunications, and biomedical implant design.

36 MATERIALS SCIENCE↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗