Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse data representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics

Large language models have revolutionized artificial intelligence by enabling large, generalizable models trained through self-supervision. This paradigm has inspired the development of scientific foundation models (FMs). However, applying this capability to experimental particle physics is challenging due to the sparse, spatially distributed nature of detector data, which differs dramatically from natural language. This work addresses if an FM for particle physics can scale and generalize across diverse tasks. We introduce a new dataset with more than 11 million particle collision events and a suite of downstream tasks and labeled data for evaluation. We propose a novel self-supervised training method for detector data and demonstrate its neural scalability with models that feature up to 188 million parameters. With frozen weights and task-specific adapters, this FM consistently outperforms baseline models across all downstream tasks. The performance also exhibits robust data-efficient adaptation. Further analysis reveals that the representations extracted by the FM are task-agnostic but can be specialized via a single linear mapping for different downstream tasks.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

AIVT: Inference of turbulent thermal convection from measured 3D velocity data by physics-informed Kolmogorov-Arnold networks

We propose the artificial intelligence velocimetry-thermometry (AIVT) method to reconstruct a continuous and differentiable representation of the temperature and velocity in turbulent convection from measured three-dimensional (3D) velocity data. AIVT is based on physics-informed Kolmogorov-Arnold networks and trained by optimizing a loss function that minimizes residuals of the velocity data, boundary conditions, and governing equations. We apply AIVT to a set of simultaneously measured 3D temperature and velocity data of Rayleigh-Bénard convection, obtained by combining particle image thermometry and Lagrangian particle tracking. This enables us to directly compare machine learning results to true volumetric, simultaneous temperature and velocity measurements. We demonstrate that AIVT can reconstruct and infer continuous, instantaneous velocity and temperature fields and their gradients from sparse experimental data at a high resolution, providing an additional approach for understanding thermal turbulence.

Science & Technology - Other Topics↗

Data Mining and Optimization Tools for Developing Engine Parameters Tools

This project was awarded for understanding the problem and developing a plan for Data Mining tools for use in designing and implementing an Engine Condition Monitoring System. From the total budget of $5,000, Tricia and I studied the problem domain for developing ail Engine Condition Monitoring system using the sparse and non-standardized datasets to be available through a consortium at NASA Lewis Research Center. We visited NASA three times to discuss additional issues related to dataset which was not made available to us. We discussed and developed a general framework of data mining and optimization tools to extract useful information from sparse and non-standard datasets. These discussions lead to the training of Tricia Erhardt to develop Genetic Algorithm based search programs which were written in C++ and used to demonstrate the capability of GA algorithm in searching an optimal solution in noisy datasets. From the study and discussion with NASA LERC personnel, we then prepared a proposal, which is being submitted to NASA for future work for the development of data mining algorithms for engine conditional monitoring. The proposed set of algorithm uses wavelet processing for creating multi-resolution pyramid of the data for GA based multi-resolution optimal search. Wavelet processing is proposed to create a coarse resolution representation of data providing two advantages in GA based search: 1. We will have less data to begin with to make search sub-spaces. 2. It will have robustness against the noise because at every level of wavelet based decomposition, we will be decomposing the signal into low pass and high pass filters.

Dhawan, Atam P.↗

Graph theory inspired anomaly detection at the LHC

Designing model-independent anomaly detection algorithms for analyzing LHC data remains a central challenge in the search for new physics, due to the high dimensionality of collider events. In this work, we develop a graph autoencoder as an unsupervised, model-agnostic tool for anomaly detection, using the LHC Olympics dataset as a benchmark. By representing jet constituents as a graph, we introduce a method to systematically control the information available to the model through sparse graph constructions that serve as physically motivated inductive biases. Specifically, (1) we construct graph autoencoders based on locally rigid Laman graphs and globally rigid unique graphs, and (2) we explore the clustering of jet constituents into subjets to interpolate between high- and low-level input representations. We obtain the best performance, measured in terms of the Significance Improvement Characteristic curve for an intermediate level of subjet clustering and certain sparse unique graph constructions. We further investigate the role of graph connectivity in jet classification tasks. Our results demonstrate the potential of leveraging graph-theoretic insights to refine and increase the interpretability of machine learning tools for collider experiments.

Automation↗

Dynamic global vegetation models underestimate net CO 2 flux mean and inter-annual variability in dryland ecosystems

Despite their sparse vegetation, dryland regions exert a huge influence over global biogeochemical cycles because they cover more than 40% of the world surface (Schimel 2010 Science 327 418–9). It is thought that drylands dominate the inter-annual variability (IAV) and long-term trend in the global carbon (C) cycle (Poulter et al 2014 Nature 509 600–3, Ahlstrom et al 2015 Science 348 895–9, Zhang et al 2018 Glob. Change Biol. 24 3954–68). Projections of the global land C sink therefore rely on accurate representation of dryland C cycle processes; however, the dynamic global vegetation models (DGVMs) used in future projections have rarely been evaluated against dryland C flux data. Here, we carried out an evaluation of 14 DGVMs (TRENDY v7) against net ecosystem exchange (NEE) data from 12 dryland flux sites in the southwestern US encompassing a range of ecosystem types (forests, shrub- and grasslands). We find that all the models underestimate both mean annual C uptake/release as well as the magnitude of NEE IAV, suggesting that improvements in representing dryland regions may improve global C cycle projections. Across all models, the sensitivity and timing of ecosystem C uptake to plant available moisture was at fault. Spring biases in gross primary production (GPP) dominate the underestimate of mean annual NEE, whereas models' lack of GPP response to water availability in both spring and summer monsoon are responsible for inability to capture NEE IAV. Errors in GPP moisture sensitivity at high elevation forested sites were more prominent during the spring, while errors at the low elevation shrub and grass-dominated sites were more important during the monsoon. We propose a range of hypotheses for why model GPP does not respond sufficiently to changing water availability that can serve as a guide for future dryland DGVM developments. Our analysis suggests that improvements in modeling C cycle processes across more than a quarter of the Earth's land surface could be achieved by addressing the moisture sensitivity of dryland C uptake.

54 ENVIRONMENTAL SCIENCES↗

Design of Digital Twin Sensing Strategies Via Predictive Modeling and Interpretable Machine Learning

This work develops a methodology for sensor placement and dynamic sensor scheduling decisions for digital twins. The digital twin data assimilation is posed as a classification problem, and predictive models are used to train optimal classification trees that represent the map from observed data to estimated digital twin states. In addition to providing a rapid digital twin updating capability, the resulting classification trees yield an interpretable mathematical representation that can be queried to inform sensor placement and sensor scheduling decisions. The proposed approach is demonstrated for a structural digital twin of a 12 ft wingspan unmanned aerial vehicle. Offline, training data are generated by simulating scenarios using predictive reduced-order models of the vehicle in a range of structural states. Furthermore, these training data can be further augmented using experimental or other historical data. In operation, the trained classifier is applied to observational data from the physical vehicle, enabling rapid adaptation of the digital twin in response to changes in structural health. Within this context, we study the performance of the optimal tree classifiers and demonstrate how they enable explainable structural assessments from sparse sensor measurements and also inform optimal sensor placement.

47 OTHER INSTRUMENTATION↗

Bayesian differential programming for robust systems identification under uncertainty

This paper presents a machine learning framework for Bayesian systems identification from noisy, sparse and irregular observations of nonlinear dynamical systems. The proposed method takes advantage of recent developments in differentiable programming to propagate gradient information through ordinary differential equation solvers and perform Bayesian inference with respect to unknown model parameters using Hamiltonian Monte Carlo sampling. This allows an efficient inference of the posterior distributions over plausible models with quantified uncertainty, while the use of sparsity-promoting priors enables the discovery of interpretable and parsimonious representations for the underlying latent dynamics. A series of numerical studies is presented to demonstrate the effectiveness of the proposed methods, including nonlinear oscillators, predator–prey systems and examples from systems biology. Taken together, our findings put forth a flexible and robust workflow for data-driven model discovery under uncertainty. All codes and data accompanying this article are available at https://bit.ly/34FOJMj .

Science & Technology - Other Topics↗

Global field reconstruction from sparse sensors with Veronoi tessellation-assisted deep learning

Achieving accurate and robust global situational awareness of a complex time-evolving field from a limited number of sensors has been a longstanding challenge. This reconstruction problem is especially difficult when sensors are sparsely positioned in a seemingly random or unorganized manner, which is often encountered in a range of scientific and engineering problems. Moreover, these sensors can be in motion and can become online or offline over time. The key leverage in addressing this scientific issue is the wealth of data accumulated from the sensors. As a solution to this problem, we propose a data-driven spatial field recovery technique founded on a structured grid-based deep-learning approach for arbitrary positioned sensors of any numbers. It should be noted that the naïve use of machine learning becomes prohibitively expensive for global field reconstruction and is furthermore not adaptable to an arbitrary number of sensors. In the present work, we consider the use of Voronoi tessellation to obtain a structured-grid representation from sensor locations enabling the computationally tractable use of convolutional neural networks. One of the central features of the present method is its compatibility with deep-learning based super-resolution reconstruction techniques for structured sensor data that are established for image processing. The proposed reconstruction technique is demonstrated for unsteady wake flow, geophysical data, and three-dimensional turbulence. The current framework is able to handle an arbitrary number of moving sensors, and thereby overcomes a major limitation with existing reconstruction methods. The presented technique opens a new pathway towards the practical use of neural networks for real-time global field estimation.

Fukami, Kai↗

Global Assimilation of Loon Stratospheric Balloon Observations

Project Loon has an overall goal of providing worldwide internet coverage using a network of long-durationsuper-pressure balloons. Since 2013, Loon has launched over 1600 balloons from multiple tropical and middlelatitude locations. These GPS tracked balloon trajectories provide lower stratospheric wind information overthe oceans and remote land areas where traditional radiosonde soundings are sparse, thus providing uniquecoverage of lower stratospheric winds. To fully investigate these Loon winds we: 1) compare the Loon windsto winds produced by a global data assimilation system (DAS: NASA GEOS) and 2) assimilate the Loon windsinto the same comprehensive DAS. Results show that in middle latitudes the Loon winds and DAS winds agreewell, and the Loon wind assimilation has only a minor impact on the forecasts. However, in the Tropics, thereis often a substantial difference between the assimilated winds and the observed Loon winds, of 8 m/s or morein magnitude. In these cases, assimilating the Loon winds significantly improves the meteorological analysesand subsequently the forecasts of the Loon winds. By highlighting cases where the Loon and DAS winds differ,these results can lead to improved understanding of stratospheric winds, especially in the tropics, as well asaiding analyses of the representation of dynamical forcing mechanisms in the GEOS model.

Coy, Lawrence↗

Robust Group Subspace Recovery: A New Approach for Multi-Modality Data Fusion

Robust Subspace Recovery (RoSuRe) algorithm was recently introduced as a principled and numerically efficient algorithm that unfolds underlying Unions of Subspaces (UoS) structure, present in the data. The union of Subspaces (UoS) is capable of identifying more complex trends in data sets than simple linear models. In this work, we build on and extend RoSuRe to prospect the structure of different data modalities individually. We propose a novel multi-modal data fusion approach based on group sparsity which we refer to as Robust Group Subspace Recovery (RoGSuRe). Relying on a bi-sparsity pursuit paradigm and non-smooth optimization techniques, the introduced framework learns a new joint representation of the time series from different data modalities, respecting an underlying UoS model. We subsequently integrate the obtained structures to form a unified subspace structure. The proposed approach exploits the structural dependencies between the different modalities data to cluster the associated target objects. The resulting fusion of the unlabeled sensors’ data from experiments on audio and magnetic data has shown that our method is competitive with other state of the art subspace clustering methods. The resulting UoS structure is employed to classify newly observed data points, highlighting the abstraction capacity of the proposed method.

47 OTHER INSTRUMENTATION↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Global Assimilation of EOS-Aura Data as a Means of Mapping Ozone Distribution in the Lower Stratosphere and Troposphere

Ozone in the lower stratosphere and the troposphere plays an important role in forcing the climate. However, the global ozone distribution in this region is not well known because of the sparse distribution of in-situ data and the poor sensitivity of satellite based observations to the lowermost of the atmosphere. The Ozone Monitoring Instrument (OMI) and Microwave Limb Sounder (MLS) instruments on EOS-Aura provide information on the total ozone column and the stratospheric ozone profile. This data has been assimilated into NASA s Global Earth Observing System, Version 5 (GEOS-5) data assimilation system (DAS). We will discuss the results of assimilating three years of OMI and MLS data into GEOS-5. This data was assimilated alongside meteorological observations from both conventional sources and satellite instruments. Previous studies have shown that combining observations from these instruments through the Trajectory Tropospheric Ozone Residual methodology (TTOR) or using data assimilation can yield useful, yet low biased, estimates of the tropospheric ozone budget. We show that the assimilated ozone fields in this updated version of GEOS-5 exhibit an excellent agreement with ozone sonde and High Resolution Dynamics Limb Sounder (HIRDLS) data in the lower stratosphere in terms of spatial and temporal variability as well as integrated ozone abundances. Good representation of small-scale vertical features follows from combining the MLS data with the assimilated meteorological fields. We then demonstrate how this information can be used to calculate the Stratosphere - Troposphere Exchange of ozone and its contribution to the tropospheric ozone column in GEOS-5. Evaluations of tropospheric ozone distributions from the assimilation will be made by comparisons with sonde and other in-situ observations.

Wargan, Krzysztof↗

Implementing and Benchmarking the Locally Competitive Algorithm on the Loihi 2 Neuromorphic Processor

Neuromorphic processors have garnered considerable interest in recent years for their potential in enabling energy-efficient and high-speed computing. The Locally Competative Algorithm (LCA) has been utilized for power efficient sparse coding on neuromophic processors, including the first Loihi processor \cite{appletospikes, loihi1}. With the Loihi 2 processor enabling custom neuron models and graded spike communication, more complex implementations of LCA are possible \cite{loihi2}. We present a new implementation of LCA designed for the Loihi 2 processor and perform an initial set of benchmarks comparing it to LCA on CPU and GPU devices. In these experiments LCA on Loihi 2 is faster and orders of magnitude more efficient, while maintaining similar reconstruction quality. We find this performance improvement increases as the LCA parameters are tuned towards greater representation sparsity. Our study highlights the potential of neuromorphic processors, particularly Loihi 2, in enabling intelligent,autonomous, real-time processing on small robots, satellite where there are strict SWaP (small, lightweighr, and low-power) requirement. By demonstrating the superior performance of LCA on Loihi 2 compared to conventional computing device, our study suggests that Loihi 2 could be a valuable tool in advancing these types of applications. Overall, our study highlights the potential of neuromorphic processors for efficient and accurate data processing on resource-constrained devices.

Parpart, Gavin G.↗

The Influence of Prescribed Boundary Conditions on Near-Surface Temperature over the Arctic in the MERRA-2 Atmospheric Model

An accurate historical record of evolving Arctic conditions is integral to furthering our understanding of climate processes and to providing a foundation for predicting future climate scenarios in northern high latitudes. Atmospheric reanalyses are seen as an important source of information on the recent past for the data-sparse Arctic region. An assessment of near-surface Arctic air temperatures finds significant discrepancies among the various modern reanalyses. An important point is the treatment of surface boundary conditions: specifically, the sea ice cover and sea surface temperatures (SSTs) over the Arctic Ocean. Reanalyses use different methodologies and data sources for SSTs and sea ice concentration boundary forcing. Notably, the Modern Era Retrospective analysis for Research and Applications, version 2 (MERRA-2) and the European Centre for Medium-Range Weather Forecasts Interim Re-Analysis (ERA-Interim) both use boundary forcing derived from the Operational Sea Surface Temperature and Sea Ice Analysis (OSTIA) over an extended, overlapping period of time. This allows for an examination of differences between the two systems while both concurrently employ the same fractional sea ice coverage. To further understand these differences, an ensemble of AMIP-style simulations using the MERRA-2 atmospheric model - but without data assimilation - shows considerable differences in Arctic temperatures as compared to reanalyses, particularly in autumn and winter months. Results from the AMIP simulations suggest that the surface representation over sea ice used in the MERRA-2 model provides an intrinsic warm bias and obfuscates Arctic Amplification, an established feature present in observations and reanalyses. An additional ensemble of AMIP-style simulations using the MERRA-2 atmospheric model was performed using boundary conditions derived from the ERA-Interim reanalysis. An in-depth comparison of surface temperatures over the Arctic from the two reanalyses and two AMIP-style ensembles will be presented, along with an assessment of the effects of the varying Arctic temperature time series on the atmospheric general circulation and energy budget.

Collow, Allison B. Marquardt↗

Latent code-based fusion: A Volterra neural network approach

We propose a deep structure encoder using Volterra Neural Networks (VNNs) to seek a latent representation of multi-modal data whose features are jointly captured by a union of subspaces. The so-called self-representation embedding of the latent codes leads to a simplified fusion which is driven by a similarly constructed decoding. The Volterra Filter architecture achieved reduction in parameter complexity is primarily due to controlled non-linearities being introduced by the higher-order convolutions in lieu of generalized activation functions. Experimental results on two different datasets have shown a significant improvement in the clustering performance for VNNs auto-encoder over conventional Convolutional Neural Networks (CNNs) auto-encoder. In addition, we also show that the proposed approach demonstrates a much-improved sample complexity over CNN-based auto-encoder with a robust classification performance.

97 MATHEMATICS AND COMPUTING↗

Inference of phase field fracture models

The phase field approach to modeling fracture uses a diffuse damage field to represent cracks. This representation mollifies singularities that arise in computations with sharp interface models and some of the resultant difficulties in the mathematical and numerical treatment of fracture. Phase field fracture models have proven effective at representing crack propagation, branching, and merging. Specific formulations, beginning with brittle fracture, have also been shown to converge to classical solutions. Extensions to cover the range of material failure, including ductile and cohesive fracture, lead to an array of possible models. There exists a large body of literature focusing on this class of models and on the impact of model form on the predicted crack evolution. However, there have not been systematic studies into how optimal models may be chosen. Here, we take a first step in this direction by developing formal methods for identification of the best parsimonious model of phase field fracture given full-field data on the damage and deformation fields. We consider some of the main models that have been used for the degradation of elastic response due to damage and its propagation. Our approach builds upon Variational System Identification (VSI), a weak form variant of the Sparse Identification of Nonlinear Dynamics (SINDy). Furthermore, in this first communication we focus on synthetically generated data but we also consider central issues associated with the use of experimental full-field data, such as data sparsity and noise.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Remote Sensing of Lineage Functional Types for Modeling and Monitoring Biodiversity

Hyperspectral remote sensing has the potential to continuously scale plant function and plant diversity information from landscape to global extents. Numerous studies have indicated that VSWIR (400-2500 nm) reflectance properties of vegetation capture evolutionarily conserved biochemical, structural, and other functional attributes of plant species. Spectral properties conserved in plants provide the opportunity to both 1) aggregate species into lineages with improved classification accuracy and 2) link those lineages directly to plant traits. Full realization of this goal will enable parameterization of Land Surface Models (LSMs) with remotely sensed information, e.g., canopy nitrogen, and better representations of biodiversity and functional diversity in biogeographic studies. In this study, we use hyperspectral AVIRIS data from the 2013 HyspIRI campaign over the Southern Sierra Nevada, California flight box to investigate the potential for incorporating evolutionary thinking into landcover classification. We link the airborne hyperspectral data with vegetation plot data from roughly 1372 surveys and a phylogeny representing 1361 species. We aggregate species into lineages ranging from species level groups down to similar number of Plant Functional Types as often used in LSMs. We assessed the ability of Random Forest and Partial Least Squares Discriminant Analysis to discriminate across these different phylogenetic scales and determine the optimal number of lineages to classify. Although there are some temporal and spatial differences in our training data, our best approaches achieved moderate classification accuracy (Kappa > 0.65). Given an optimal number of lineages, we explored approaches to improve classifications including machine learning and unmixing approaches. This work suggests that lineage-based methods may be a promising way to leverage the huge amounts of data that will come from high resolution and high return interval hyperspectral data planned for the Surface Biology and Geology mission with sparsely sampled existing ground-based ecological data.

Hyperspectral↗