Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Vision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images

Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of 12 state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of 469K question8 answer pairs involving 30K images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images

Maruf, M [Virginia Tech, Blacksburg]

SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical Segmentation

This paper presents a novel approach to addressing the long-sequence problem in high-resolution medical images for Vision Transformers (ViTs). Using smaller patches as tokens can enhance ViT performance, but quadratically increases computation and memory requirements. Therefore, the common practice for applying ViTs to high-resolution images is either to: (a) employ complex sub-quadratic attention schemes or (b) use large to medium-sized patches and rely on additional mechanisms within the model to capture the spatial hierarchy of details. We propose Symmetrical Hierarchical Forest (SHF), a lightweight approach that adaptively patches the input image to increase token information density and encode hierarchical spatial structures into the input embedding. We then apply a reverse depatching scheme to the output embeddings of the transformer encoder, eliminating the need for convolution-based decoders. Unlike previous methods that modify attention mechanisms or use a complex hierarchy of interacting models, SHF can be retrofitted to any ViT model to allow it to learn the hierarchical structure of details in high-resolution images without requiring architectural changes. Experimental results demonstrate significant gains in computational efficiency and performance: on the PAIP WSI dataset, we achieved a 3∼32×speedup or a 2.95%∼7.03% increase in accuracy (measured by Dice score) at a 64K2 resolution with the same computational budget, compared to state-of-the-art production models. On the 3D medical datasets BTCV and KiTS, training was 6×faster, with accuracy gains of 6.93% and 5.9%, respectively, compared to models without SHF.

Zhang, Enzhi [Hokkaido University, Japan]

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja

Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization

Deploying geospatial foundation models on resource-constrained edge devices demands compact architectures that maintain high downstream performance. However, their large parameter counts and the accuracy loss often induced by compression limit practical adoption.In this work, we leverage manifold-constrained optimization framework DLRT to compress large vision transformer–based geospatial foundation models during transfer learning. By enforcing structured low-dimensional parameterizations aligned with downstream objectives, this approach achieves strong compression while preserving task-specific accuracy. We show that the method outperforms of-the-shelf low-rank methods as LoRA. Experiments on diverse geospatial benchmarks confirm substantial parameter reduction with minimal accuracy loss, enabling high-performing, on-device geospatial models.

Snyder, Thomas [Yale University]

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for this task, yet no systematic evaluation exists of how well vision-language models (VLMs) interpret them. We introduce QCalEval, the first VLM benchmark for quantum calibration plots: 243 samples across 87 scenario types from 22 experiment families, spanning superconducting qubits and neutral atoms, evaluated on six question types in both zero-shot and in-context learning settings. The best general-purpose zero-shot model reaches a mean score of 72.3, and many open-weight models degrade under multi-image in-context learning, whereas frontier closed models improve substantially. A supervised fine-tuning ablation at the 9-billion-parameter scale shows that SFT improves zero-shot performance but cannot close the multimodal in-context learning gap. As a reference case study, we release NVIDIA Ising Calibration 1, an open-weight model based on Qwen3.5-35B-A3B that reaches 74.7 zero-shot average score.

Cao, Shuxiang

CTRL-STEER: Closed-Loop Neuron Activation Control in Vision-Language-Action Models

Vision-Language-Action (VLA) models enable test-time behavioral steering via neuron-level interventions, but existing methods use fixed strengths and operate in open loop. This static modulation fails under evolving task dynamics, leading to overcorrection, oscillations, and reduced task success—especially for temporal attributes like speed. We propose CTRL-STEER, a control-theoretic framework that casts activation steering as closed-loop feedback with adaptive, time-varying interventions. Instead of assuming neurons encode temporal concepts, we steer along motion-aligned residual directions and regulate intervention magnitude via feedback. We instantiate this with both PID and reinforcement learning controllers that jointly optimize concept adherence and task success. Experiments on fine-tuned OpenVLA policies across four LIBERO suites show improved stability and a better steering–success trade-off over fixed-coefficient baselines, without retraining the base model.

Babu, Abhijith [Florida International University,

Describing Point Defect Topology in 2D Energy Materials through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. Here we employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science.

2d materials

Computer Vision on Edge Devices for the Short Term Prediction of Cloud Cover

Edge Computing and IoT are important pieces of today's technological landscape. Here, we build a low-cost IoT sensor for sky imaging and program it using AWS GreenGrass, one of the leading IoT platforms. We demonstrate remote reprogramming of this device to load software that predicts sun shading events through the linear advection method, which is a baseline algorithm that can be used to benchmark algorithmic improvements in future work. Some future directions for sky imaging research are enumerated.

14 SOLAR ENERGY

Describing Point Defect Topology in 2D Energy Materials Through Computer Vision

Point defects such as vacancies and impurity atoms strongly impact the performance of 2D materials. Traditional efforts often rely on manual detection, a process that is time-intensive, prone to human error, and challenging to scale. Here we leverage machine learning (ML) methods to identify and quantify vacancies within 2D transition metal carbides (Ti3C2, MXenes), aiming to expedite detection while improving accuracy. MXenes exhibit valuable defect-defined electrochemical properties, but we currently lack statistical understanding of defect topology needed to fully harness these materials. We employ a convolutional neural network for semantic segmentation of experimental MXene images, opening an opportunity to conduct a rigorous statistical study on defect hierarchy while investigating local relaxation in the lattice. We show how the integration of ML can yield fundamental insight into point defects, providing a powerful tool that will play an increasingly crucial role in the future of materials science.

2D materials

Biorefinery siting and sizing to achieve the US Billion‐Ton Bioeconomy vision: A case study using a gasification–Fischer–Tropsch process

Achieving a secure, abundant, and affordable energy future requires a robust and adaptable energy strategy, with bioenergy playing a pivotal role. Biomass-based energy presents a promising pathway to use domestic resources while fostering economic opportunities in rural areas. Despite the potential to source more than 1 billion dry short tons of biomass annually in the US, significant infrastructure and economic barriers hinder full utilization for energy production. This study used the Biofuel Infrastructure, Logistics, and Transportation (BILT) model to assess biorefinery siting and scale and determine the number and size of facilities required to maximize use of the US biomass potential. A spatially agnostic approach first assessed the effects of facility capacity and transportation constraints on biomass use. Then, a spatially explicit analysis integrated county-level biomass availability from the US Department of Energy's 2023 Billion-Ton Report and technoeconomic assessments to evaluate different biorefinery deployment scenarios. The results indicate that an optimized mix of facility sizes is essential to leverage biomass resources fully across varying regional production densities to maximize use of the US biomass potential. Larger biorefineries or co-located smaller facilities significantly enhance biomass use while reducing costs through economies of scale. These findings underscore the importance of strategically balancing facility capacity and spatial distribution to optimize the bioenergy supply chain. In conclusion, this study provides critical insights for advancing the US bioenergy economy by aligning biorefinery deployment with biomass resource availability and economic viability.

BILT Model

Soybean genomics research community strategic plan: A vision for 2024–2028

Abstract This strategic plan summarizes the major accomplishments achieved in the last quinquennial by the soybean [Glycine max(L.) Merr.] genetics and genomics research community and outlines key priorities for the next 5 years (2024–2028). This work is the result of deliberations among over 50 soybean researchers during a 2‐day workshop in St Louis, MO, USA, at the end of 2022. The plan is divided into seven traditional areas/disciplines: Breeding, Biotic Interactions, Physiology and Abiotic Stress, Functional Genomics, Biotechnology, Genomic Resources and Datasets, and Computational Resources. One additional section was added, Training the Next Generation of Soybean Researchers, when it was identified as a pressing issue during the workshop. This installment of the soybean genomics strategic plan provides a snapshot of recent progress while looking at future goals that will improve resources and enable innovation among the community of basic and applied soybean researchers. We hope that this work will inform our community and increase support for soybean research.

Genetics & Heredity

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE

A U.S. Scientific Community Vision for Sustained Earth Observations of Greenhouse Gases to Support Local to Global Action

Managing carbon stocks in the land, ocean, and atmosphere under changing climate requires a globally‐integrated view of carbon cycle processes at local and regional scales. The growing Earth Observation (EO) record is the backbone of this multi‐scale system, providing local information with discrete coverage from surface measurements and regional information at global scale from satellites. Carbon flux information, anchored by inverse estimates from spaceborne Greenhouse Gas (GHG) concentrations, provides an important top‐down view of carbon emissions and sinks, but currently lacks global continuity at assessment and management scales (<100 km). Partial‐column data can help separate signals in the boundary layer from the overlying atmosphere, providing an opportunity to enhance surface sensitivity and bring flux resolution down from that of column‐integrated data (100–500 km). Based on a workshop held in September 2024, the carbon cycle community envisions a carbon observation system leveraging GHG partial columns in the lower and upper troposphere to weave together information across scales from surface and satellite EO data, and integration of top‐down/bottom‐up analyses to link process understanding to global assessment.

Parazoo, Nicholas C. [California Institute of Tech