Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Uncertainty analysis of correlated parameters in automated reaction mechanism generation

Abstract Uncertainty analysis is a useful tool for inspecting and improving detailed kinetic mechanisms because it can identify the greatest sources of model output error. Owing to the very nonlinear relationship between kinetic and thermodynamic parameters and computed concentrations, model predictions can be extremely sensitive to uncertainties in some parameters while uncertainties in other parameters can be irrelevant. Error propagation becomes even more convoluted in automatically generated kinetic models, where input uncertainties are correlated through kinetic rate rules and thermodynamic group values. Local and global uncertainty analyses were implemented and used to analyze error propagation in Reaction Mechanism Generator (RMG), an open‐source software for generating kinetic models. A framework for automatically assigning parameter uncertainties to estimated thermodynamics and kinetics was created, enabling tracking of correlated uncertainties. Local first‐order uncertainty propagation was implemented using sensitivities computed natively within RMG. Global uncertainty analysis was implemented using adaptive Smolyak pseudospectral approximations as implemented in the MIT Uncertainty Quantification Library to efficiently compute and construct polynomial chaos expansions to approximate the dependence of outputs on a subset of uncertain inputs. Cantera was used as a backend for simulating the reactor system in the global analysis. Analyses were performed for a phenyldodecane pyrolysis model. Local and global methods demonstrated similar trends; however, many uncertainties were significantly overestimated by the local analysis. Both local and global analyses show that correlated uncertainties based on kinetic rate rules and thermochemical groups drastically reduce a model's degrees of freedom and have a large impact on the determination of the most influential input parameters. These results highlight the necessity of incorporating uncertainty analysis in the mechanism generation workflow.

Gao, Connie W.↗

Advances in Multimodal Characterization of Structural Materials

The myriad detectors and instruments now available for materials characterization provide researchers with an ever-growing suite of tools to probe material behavior. Progress in the development of instrumentation and workflows that enable the collection, and leverage the potential, of various data modalities have provided novel insights into material behavior. Using data across multiple length scales, or performing complementary analyses of in situ and ex situ data, can help reveal a more complete picture of dynamic processes or material structure. However, the accurate combination, or fusion, of these disparate data modalities presents new challenges. Differences in resolution, as well as the varying length scales at which physical phenomena are exploited to generate these data, necessitate novel approaches to accurately interpret and combine these data. Furthermore, the papers within this special topic focus on the collection and fusion of multimodal data to better understand structural materials. From new frameworks and workflows for data segmentation and analysis, process monitoring, enhancing simulations, or interrogating mechanical response, these papers reveal the potential benefits of utilizing multimodal data.

36 MATERIALS SCIENCE↗

Assessing pore network heterogeneity across multiple scales to inform CO2 injection models

Geologic heterogeneity is a key feature that must be considered when translations of scaled data are performed. This paper presents the assessment of geologic heterogeneity using a multiscale workflow that includes image analysis-based methods coupled with well log analysis to provide data in which fractals and machine learning methods estimate the carbon dioxide (CO 2 ) storage resource potential of a reservoir. The heterogeneity of rock properties of the complex Bell Creek reservoir in Montana, USA, was explored at the pore scale (~nm to mm), core scale (~mm to m), and well scale (~cm to m). The data used in this study included advanced image analysis of micro-CT (computed tomography) images (pore scale), thin sections (pore scale), plugs and core images (core scale) and well logs (well scale). The micro-CT images were segmented using a U-net segmentation approach into objects of pores and grains. Further, the segmented images were reconstructed into subvolumes of different sizes. Physical properties (porosity and permeability) and fractal dimensions were calculated for the various subvolumes, and Lorenz coefficient (Lc) values, a single parameter to describe the degree of heterogeneity within a pay zone section, were calculated from thin-section images and well logs. Porosity and fractal dimension values were used to estimate the 188-µm threshold of representative elementary volume (REV) in this study. Both the Lc and fractal dimension values were found to be negatively correlated. When these two parameters are combined, it is possible to discern differences in the complex porous networks of the samples analyzed in this study.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)

This data release provides all data and code used in the paper " "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)" to model stream temperature, evaluate, and assess results. The associated manuscript explores current open questions in prediction in ungauged and unmonitored basins concerning top-down versus bottom-up approaches, tradeoffs between data available and input requirements, and the appropriate representation of catchment attributes as inputs to deep learning models. Modeling was done primarily with long short-term memory (LSTM) models, and stream site coverage spans 1362 locations across the conterminous United States. The data is organized into these items items:Code repository and data for the paper " "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)".Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code: - data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- error_analysis_attribute_and_groundwater_dir.zip - workflows for the extended error analysis by stream attribute and groundwater influenceData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2024streamdata, author = {Jared Willard and Fabio Ciulla and Helen Weierbach and Vipin Kumar and Charuleka Varadharajan}, title = {Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models"}, year = {2024}, doi = {10.15485/2448016}, publisher = {ESS-DIVE Repository}, url = {https://doi.org/10.15485/2448016}}MLA: Willard, Jared, et al. Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models". 2024. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Predicting Dynamic-to-Static Correction Factor from Petrophysical Data and Chemostratigraphy using Unsupervised Machine Learning

Estimating static mechanical properties of stratigraphic layers is critical for optimizing subsurface engineering applications. To estimate dynamic-to-static correction factor F ds (static-to-dynamic Young’s modulus ratio) across the Caney shale interval in Oklahoma, USA, we integrated triaxial test measurements and petrophysical data, including well logs and X-ray fluorescence (XRF) using unsupervised machine learning (ML). We used a novel workflow that includes principal component analysis (PCA) to reduce data set dimensionality of well logs and XRF data sets—both separately and combined—creating three scenarios, and later applied inverse distance weighting (IDW) to derive F ds profiles for these scenarios. Furthermore, we applied K-means clustering on each scenario to predict depositional facies, and built a stiffness zonation profile through chemostratigraphic analysis of the terrigenous elements to validate the predicted F ds . The predicted F ds profile from each scenario using the PCA-IDW method was compared with the constant F ds approach from our previous study by calculating the root mean square error (RMSE). The combined data sets scenario yielded the lowest RMSE value of 0.113, while the RMSE values for the well logs and XRF scenarios were 0.131 and 0.129, respectively. In addition, the predicted F ds from the XRF scenario well-matched the stiffness zonation from the chemostratigraphic analysis that was built using the optimized K-means clustering of nine clusters for that scenario. These methods and findings offer a valuable tool for refining lithological classification and improving the F ds profile, potentially enhancing drilling and stimulation strategies for subsurface energy engineering applications.

clastic rock↗

A Generative Model for Synthetic Electroluminescence Images

This work will develop modular, open-source model and analysis components including crack detection workflow and parameterization for quantitative inspection of large EL large datasets. These tools will allow users to quickly and accurately assess the extent and types of cracking in their modules. Measured statistical distributions of crack parameters, together with the imposed stress and electrical properties will be used to generate models to predict future crack behavior and power loss.

Pierce, Benjamin Garrett↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Provenance in Data Interoperability for Multi-Sensor Intercomparison

As our inventory of Earth science data sets grows, the ability to compare, merge and fuse multiple datasets grows in importance. This requires a deeper data interoperability than we have now. Efforts such as Open Geospatial Consortium and OPeNDAP (Open-source Project for a Network Data Access Protocol) have broken down format barriers to interoperability; the next challenge is the semantic aspects of the data. Consider the issues when satellite data are merged, cross-calibrated, validated, inter-compared and fused. We must match up data sets that are related, yet different in significant ways: the phenomenon being measured, measurement technique, location in space-time or quality of the measurements. If subtle distinctions between similar measurements are not clear to the user, results can be meaningless or lead to an incorrect interpretation of the data. Most of these distinctions trace to how the data came to be: sensors, processing and quality assessment. For example, monthly averages of satellite-based aerosol measurements often show significant discrepancies, which might be due to differences in spatio- temporal aggregation, sampling issues, sensor biases, algorithm differences or calibration issues. Provenance information must be captured in a semantic framework that allows data inter-use tools to incorporate it and aid in the intervention of comparison or merged products. Semantic web technology allows us to encode our knowledge of measurement characteristics, phenomena measured, space-time representation, and data quality attributes in a well-structured, machine-readable ontology and rulesets. An analysis tool can use this knowledge to show users the provenance-related distrintions between two variables, advising on options for further data processing and analysis. An additional problem for workflows distributed across heterogeneous systems is retrieval and transport of provenance. Provenance may be either embedded within the data payload, or transmitted from server to client in an out-of-band mechanism. The out of band mechanism is more flexible in the richness of provenance information that can be accomodated, but it relies on a persistent framework and can be difficult for legacy clients to use. We are prototyping the embedded model, incorporating provenance within metadata objects in the data payload. Thus, it always remains with the data. The downside is a limit to the size of provenance metadata that we can include, an issue that will eventually need resolution to encompass the richness of provenance information required for daata intercomparison and merging.

Lynnes, Chris↗

Machine Learning for Well Log Analysis in Uranium Mining

This project explores the use of Artificial Intelligence (AI) and Machine Learning (ML) techniques to automate well log analysis for uranium mining. Geophysical log data—spontaneous potential, resistivity, and gamma ray—were used to classify lithology, correlate well logs and identify roll front zonation patterns, which are critical for locating uranium ore bodies. Supervised ML algorithms such as eXtreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), and Random Forest were trained to classify lithology with high accuracy. Gradient Boosting Machines (GBM), XGBoost, Random Forest, and Neural Networks were also used for role front zone identification. Moreover, a Fast Dynamic Time Warping (FastDTW) algorithm was employed for well log correlation. Additionally, sample lag was addressed using dynamic programming. Results demonstrate the potential of AI and ML to streamline well log analysis and enhance uranium exploration workflows.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Comprehensive Analysis of Streaming and Shutdown Dose Rate Experiments at JET with ORNL Fusion Neutronics Workflows

Current experimental fusion systems and conceptual designs of fusion pilot plants (FPPs) are growing in complexity and size. Several radiation metrics are crucial to the safe operation of fusion machines, including neutron flux streaming through openings and the shutdown dose rate (SDDR). Most current designs of advanced experimental fusion systems—and the most probable candidates for FPPs—are based on the tokamak concept, which is prone to neutron streaming through the myriad openings needed for diagnostic and support systems. SDDR is caused by decay gamma rays from radionuclides that become activated by neutrons during the operation of a fusion system that use deuterium-deuterium (DD), tritium-tritium, or deuterium-tritium plasma. Because computational tools have become essential for determining these radiation metrics, they must be validated against reliable and applicable experimental data. Experiments at the Joint European Torus (JET) provide a unique source of experimental data for validating computational tools and nuclear data used to determine SDDR and neutron fluxes in streaming-dominated geometries. Here, this paper presents the comprehensive analysis of the high-performance DD JET SDDR, and streaming experiments performed using Oak Ridge National Laboratory (ORNL) fusion workflows. The computational results were compared with experimental results that consist of online SDDR measurements with ionization chambers and neutron fluence streaming measurements using thermoluminescent detectors. The ratio of calculated-to-experimental SDDR values ranges from 0.6 to 2.5, and the streaming results range from 0.5 to 8.0. Future work will include analyzing the JET 2021 DTE2 campaign alongside the integration of the Shift Monte Carlo transport code into all ORNL fusion neutronics workflows.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Improving and Automating Building Model Data Exchange

There are many instances throughout a project’s lifecycle where there arises a need for quick and accurate risk assessment of building designs. For example, an unexpected design change during construction may necessitate structural engineers to perform a seismic risk assessment on analytical models of the updated building design using high fidelity structural analysis software, such as ANSYS or Abaqus. However, the efficiency of such workflows often depends upon the interoperability of architectural design software and structural analysis software. When the quality of this interoperability is lacking or even non-existent, the efficiency of virtual engineering workflows is hampered, which increases project costs. A McGraw Hill industry survey of professional users of Building Information Modeling (BIM) technologies found that there is high demand for BIM interoperability for structural analysis, but that the value/difficulty ratio is currently too low for practical use. There have been efforts by the academic community to facilitate model data exchange between the architectural design and structural analysis domains, but such solutions have not been widely adopted by industry, face technical challenges, and oftentimes are limited in applicability for users of various BIM software. Therefore, INL is developing capabilities to improve, automate, and generalize model data exchange between architectural BIM software (e.g., Revit) and structural analysis software (e.g., SAP2000, ANSYS). The goal is to help expedite and automate as much of the pre-processing step for creating analytical models in finite element analysis software as reasonably as possible. Such a "BIM-to-FEA" conversion tool should provide direct benefit to end-users through accuracy, automation, quick turn-around, and wide applicability. To generalize the application of this BIM-to-FEA conversion tool and increase its useability among the many different commercial BIM software currently used by industry, the program is being developed with the concept of openBIM. OpenBIM is the application of non-proprietary, open data standards that allow for BIM model data exchange in a format that is accessible, retainable, and useable for all users. The most widely used open, non-proprietary data exchange format for BIM is the Industry Foundation Classes (IFC) schema. IFC is developed by buildingSMART international and is ISO certified (ISO 16739-1:2018). The BIM-to-FEA conversion tool is being developed for compatibility with typical commercial building designs of steel framed structures. The tool is currently capable of importing architectural BIM data of framed building structures, recognizing and extracting the aspects of the model that are required for structural analysis, adjusting the connectivity of frame members, and finally exporting to an analytical model stored in the IFC format. The exported IFC analytical model can then be imported into various openBIM compliant software, such as SAP2000. Such capabilities have already been tested on commercial software, as shown above, and continue to be improved. Work is underway to test the conversion on various commercial BIM software, develop a user-friendly interface, incorporate the program into the broader DeepLynx data warehouse project being developed by INL, and to eventually open-source the tool for the benefit of the community. Future development of the tool envisions the ability for efficient iterative risk assessment of generative building designs, all within a workflow utilizing open-source tools. One such open-source tool will be MOOSE, an advanced finite element analysis tool developed at INL. The conversion tool will also branch out from typical commercial building designs and will aim to incorporate nuclear construction. The aim will be to convert both structural and non-structural components of nuclear facilities, such as curved concrete containment structures and piping systems, respectively.

97 MATHEMATICS AND COMPUTING↗

A Systematic Interpretation of Subsurface Proppant Concentration from Drilling Mud Returns: Case Study from Hydraulic Fracturing Test Site (HFTS-2) in Delaware Basin

The aim of this study is generation and validation of a proppant log using analysis of drilling mud returns for child wells. Proppant log provides qualitative as well as quantitative insights into spatial distribution of proppant sand particles from prior stimulation of parent wells. While the basic methodology was developed and formalized during analysis of material collected from through fracture cores at Hydraulic Fracturing Test Site in Midland Basin (HFTS – 1), the test wells at HFTS – 2 in the neighboring Delaware Basin allowed the opportunity to validate the workflow on actual mud return samples from subsurface. As a child well is being drilled, periodic mud return samples are collected at the rig site and preserved for analysis. The workflow involves systematic cleaning of the samples including various steps such as washing, drying and segregation of samples into relevant size fractions of interest (< Mesh 20) based on specifications of pumped sand during stimulation of the parent well. Clean samples are imaged using high resolution transparency scanning. Scan images are then systematically analyzed for particles of interest using computer vision techniques. Sample counts are further validated using elemental analysis of smaller sub-samples at various depths of interest. This step is necessary to isolate proppant versus other naturally occurring minerals such as sulphates and carbonates which show similar optical properties. We successfully correlated proppant distribution against the existing parent well and validated propped versus relatively un-propped zones for a child well at the test site. The advantage of testing the proppant log concept at the HFTS – 2 site is the plethora of additional diagnostic data that is available to validate our primary observations. We can correlate spatial proppant distribution against variability in stimulation response based on independent observations such as image logs, microseismic attributes as well as DAS response, all of which tend to corroborate one another. One of our significant successes was being able to describe varying degrees of impact of the parent well along the lateral length of a stimulated child well. Our workflow represents a systematic and one-of-a-kind interpretation of spatial proppant distribution while drilling child wells. This provides unique opportunities to better understand the current state of the Downloaded from http://onepetro.org/URTECONF/proceedings-pdf/21URTC/2-21URTC/D021S031R003/2477415/urtec-2021-5189-ms.pdf/1 by Carol Worster on 28 February 2022 URTeC 5189 2 reservoir being targeted including zones which are likely more drained relative to others and how the planned completion of the child well can be improved. Lastly, this log can be useful is validating optimal well spacing in relatively new fields under development.

58 GEOSCIENCES↗

Software Project Management and Measurement on the World-Wide-Web (WWW)

We briefly describe a system for forms-based, work-flow management that helps members of a software development team overcome geographical barriers to collaboration. Our system, called the Web Integrated Software Environment (WISE), is implemented as a World-Wide-Web service that allows for management and measurement of software development projects based on dynamic analysis of change activity in the workflow. WISE tracks issues in a software development process, provides informal communication between the users with different roles, supports to-do lists, and helps in software process improvement. WISE minimizes the time devoted to metrics collection and analysis by providing implicit delivery of messages between users based on the content of project documents. The use of a database in WISE is hidden from the users who view WISE as maintaining a personal 'to-do list' of tasks related to the many projects on which they may play different roles.

Callahan, John↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

In-Transit Data Transport Strategies for Coupled AI-Simulation Workflow Patterns

Coupled AI-Simulation workflows are becoming the major workloads for HPC facilities, and their increasing complexity necessitates new tools for performance analysis and prototyping of new in-situ workflows. We present SimAI-Bench, a tool designed to both prototype and evaluate these coupled workflows. In this paper, we use SimAI-Bench to benchmark the data transport performance of two common patterns on the Aurora supercomputer: a one-to-one workflow with co-located simulation and AI training instances, and a many-to-one workflow where a single AI model is trained from an ensemble of simulations. For the one-to-one pattern, our analysis shows that node-local and DragonHPC data staging strategies provide excellent performance compared Redis and Lustre file system. For the many-to-one pattern, we find that data transport becomes a dominant bottleneck as the ensemble size grows. Our evaluation reveals that file system is the optimal solution among the tested strategies for the many-to-one pattern.

Tummalapalli, Harikrishna [Argonne National Labora↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗