Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “video adaptation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

FPV Video Adaptation for UAV Collision Avoidance

First person view (FPV) technology for unmanned aerial vehicles (UAVs) provides an immersive experience for pilots and enables various personal and commercial applications such as aerial photography, drone racing, search and rescue operations, agricultural surveillance, and structural inspection. While real time video streaming from a UAV and vision-based collision avoidance strategies have been studied in literature as separate topics, in this paper we tackle collision avoidance in FPV scenarios, taking into account network delays and real time video parameters. We present a theoretical model for obstacle collisions that considers the current communication channel conditions, the real time video parameters, and the UAV's position relative to the closest obstacle. A video adaptation algorithm is then designed, using this metric, to tune the FPV video resolution, number of re-transmission attempts, and the modulation scheme to maximize the probability of avoiding collisions. This algorithm also takes into account specific latency constraints of the application. This video algorithm was evaluated in various scenarios and its ability to respond to both distances to the obstacle as well as the communication channel conditions was demonstrated. It was found that, for the considered scenarios, the performance of the proposed adaptive algorithm was, on an average, 58.63% higher than the closest non-adaptive one in terms of maximizing the probability of avoiding collision. Such collision avoidance strategies could be used to make UAV FPV applications safer and more reliable.

47 OTHER INSTRUMENTATION↗

Developing Smart Building Technology Modules to Enhance Workforce Preparedness: A Case for AI-Driven Academic and Professional Education

Smart building technologies are resources that improve building energy efficiency and resilience, reduce carbon emissions, and provide load flexibility to the grid. However, in both academic curricula and building professionals’ continuing education, there is a lack of systematic instruction on methods to integrate multiple energy systems including distributed energy resources (DER), smart building technologies, AI (Artificial Intelligence) tools and key concepts, components, and controls, including “Internet of Things” (IoT) devices. In today’s dynamic workforce, this major gap in smart building technology education prevents stakeholders from being able to attract talent with an understanding and preparation to adopt smart building technologies in building design and operations. A federally funded project included a partnership between Slipstream and Texas A&M University (TAMU) to develop a semester-long smart building curriculum for engineering college students with the ability to adapt the contents for workforce development of professionals in building services. The final product consists of 16 training videos adapted for building professionals and the public. The educational content and training materials cover the benefits of building energy systems, the latest sensor technologies and IoT devices, all with a focus on smart building technologies. The key drivers are on topics related to smart building controls (i.e., energy management information systems), smart building control platforms, cybersecurity, grid-interactive-efficient buildings (GEBs), smart building control methods, and occupant-centric control. Although not explicitly included the technologies nod to the need for AI driven technologies to prepare engineers and industry professionals to be future ready. This paper describes the project approach, provides outlines of the training materials, and identifies lessons learned in creating the content for this course. The authors suggest ways to scale the instruction of smart building concepts to empower the workforce to accelerate the adoption of smart building technologies and AI-based teaching and learning in higher education and building sector.

99 GENERAL AND MISCELLANEOUS↗

FunDiff: diffusion models over function spaces for physics-informed generative modeling

Recent advances in generative modeling-particularly diffusion models and flow matching-have been widely used for synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. To address this, we introduce FunDiff, an efficient and robust framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generates continuous functions that can be evaluated at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, demonstrating that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We further demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results indicate that our method can generate physically consistent samples with high fidelity to the target distribution, and exhibit robustness to noisy and low-resolution data.

Wang, Sifan [Yale University, New Haven, CT (Unite↗

Development and Validation of Smart Building Technology Modules for Academic and Professional Education

Slipstream leads a team developing a semester-long smart building curriculum for college students and adapting the contents into 16 training videos for building professionals and the public. The topics cover smart building technologies related content including industry trends and benefits, building systems, sensors and IoT devices, advanced building monitoring and controls, smart building control platform, methods, and applications.

99 GENERAL AND MISCELLANEOUS↗

Smart Building Technology Training Modules for Academic and Professional Education

Smart building technologies are a new suite of resources that improve building energy efficiency and resilience, reduce carbon emissions, and provide load flexibility to the grid. However, in both college curricula and building professionals’ continuing education, there is a lack of systematic instruction on smart building technologies–topics that include smart building concepts, key components, smart building controls, “Internet of Things” (IoT) devices, and how to integrate multiple energy systems including distributed energy resources (DER). This major gap in smart building education prevents stakeholders from understanding and adopting smart building technologies in building design and operations. Slipstream leads a DOE-funded project developing a semester-long smart building curriculum for college students and adapting the contents into 16 training videos for building professionals and the general public. The education and training cover the drivers and benefits of smart building technologies, key building energy systems, the latest sensor technologies and IoT devices, and focus on topics related to smart building controls (i.e., energy management information systems, smart building control platforms, cybersecurity, grid-interactive-efficient buildings (GEBs), smart building control methods, and occupant-centric control. This paper describes the project approach, provides outlines of the training materials, and identifies lessons learned in creating the content. We also suggest ways to scale the instruction of smart building concepts to empower the workforce to accelerate the adoption of smart building technologies in the real world.

99 GENERAL AND MISCELLANEOUS↗

Virtual Reality for Shoot/No-Shoot Decision Training in Law Enforcement: A Literature Review and Research Agenda

Virtual reality (VR) can materially improve “shoot / no-shoot” (SNS) training by giving officers realistic, repeatable practice making high-stakes decisions under pressure. Traditional tools—live-fire ranges and video simulators—build basics, but they cannot adapt to each officer in real time or fully mirror the complexity of the field. VR closes that gap by creating immersive scenarios that are safer, more flexible, easier to scale across units, and able to capture objective performance data. SNS decisions are not just about marksmanship; they rely on perception, judgment, memory, and the ability to hold fire when a threat is uncertain. Effective training therefore needs realism, decision complexity, and branching outcomes that reflect the true consequences of choices. These elements strengthen recognition of hostile intent while reducing false positives and building the self-control required in ambiguous situations. VR brings specific advantages: dynamic environments, full-body interaction, and the ability to measure performance with precision—enabling targeted feedback and better transfer of learning to the street. At the same time, responsible deployment must address scenario quality (credible environments and behaviors), lawful decision models, and user wellbeing (appropriate stress levels, comfort, and safety). Sandia’s VIPER Lab is positioned to lead this work. The team combines human-performance science, AI/ML, and VR/AR development with a deep equipment bench (e.g., omnidirectional treadmill, eye-tracking, haptics, multiple HMDs). This ecosystem supports building and validating next-generation SNS training that is immersive, measurable, and trustworthy. Bottom line: Investment in VR-enabled SNS training that blends evidence-based design with careful validation and legal safeguards is expected to pay off in safer, more consistent decision-making and improved community trust, delivered through training that is practical to deploy at scale.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Learning the Temporal Effect in Infrared Thermal Videos With Long Short-Term Memory for Quality Prediction in Resistance Spot Welding

With the advances of sensing technology, in-situ infrared thermal videos can be collected from Resistance Spot Welding (RSW) processes. Each video records the formulation process of a weld nugget. The nugget evolution creates a “temporal effect” across the frames, which can be leveraged for real-time, nondestructive evaluation (NDE) of the weld quality. Currently, quality prediction with imaging data mainly focuses on optical feature extraction with Convolutional Neural Network (CNN) but does not make the most of such temporal effect. In this study, pixels corresponding to critical locations on the weld nugget surface are extracted from a video to form multivariate time series (MTS). Multivariate Adaptive Regression Splines (MARS) is used in MTS processing to remove noisy signals related to uninformative frames. A Stacked Long Short-Term Memory (LSTM) model is developed to learn from the processed MTS and then predicts weld nugget size and thickness in real-time NDE. Results from a case study on RSW of Boron steel demonstrates the improvement in prediction accuracy and computational time with the proposed method, as compared to CNN-based weld quality prediction.

Guo, Shenghan↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Producing two-dimensional dust clouds and clusters using a movable electrode for complex plasma and fundamental physics experiments

We report a Bidirectional Electrode Control Arm Assembly (BECAA) for precisely manipulating dust clouds levitated above the powered electrode in RF plasmas. The reported techniques allow the creation of perfectly 2D dust layers by eliminating off-plane particles by moving the electrode from outside the plasma chamber without altering the plasma conditions. Here, the tilting and moving of electrodes using BECAA also allows the precise and repeatable elimination of dust particles one by one to achieve any desired number of grains N without trial and error. Simultaneously acquired top and side view images of dust clusters show that they are perfectly planar or 2D. A demonstration of clusters with N = 1–28 without changing the plasma conditions is presented to show the utility of BECAA for complex plasma and statistical physics experimental design. Demonstration videos and 3D printable part files are available for easy reproduction and adaptation of this new method to repeatably produce 2D clusters in existing RF plasma chambers.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Using Machine Learning to Track Objects Across Cameras

Video surveillance is one of the most important technologies used by the International Atomic Energy Agency in international safeguards. At large, complicated facilities, multiple surveillance cameras are deployed to monitor the transfer of safeguards-relevant objects across the site. During inspections, all surveillance videos are reviewed to ensure the objects are not manipulated or diverted during transfer, a laborious, time-consuming task. This work describes using deep machine learning algorithms to track objects automatically across multiple cameras, greatly improving the efficiency of the review process. The fundamental problem in this object tracking task across multiple cameras is how to associate the same object, which may show extreme intra-class variations, such as viewpoints, occlusions, and various scales, in different and even non-overlapped cameras. Object re-identification (Re-ID) in nuclear facility video surveillance is even more challenging than classic person or vehicle Re-ID problems because different instances in the same category may display an identical appearance. One observation from nuclear facility surveillance videos is that all objects must be carted (e.g., via forklift) to move. Therefore, the spatial context information of an object, which provides the feature from the carrier, is critical for the object Re-ID task. This work proposes a two-stream convolutional neural networks model that takes features of objects and their surrounding regions into account. Moreover, the custom videos usually are gleaned from different scenes from the training data, which may have extreme variations in illumination changes and/or cluttered backgrounds. Directly applying the trained model to custom videos will dramatically decrease the performance. To tackle this problem, an advanced domain adaptation technique is proposed to mitigate the gap between the data taken from different scenes. The proposed framework will track objects of interest across a nuclear complex. The resulting tracks can be used in further analyses, such as event/activity recognition, anomaly detection, etc.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Truck Platooning Performance with ADAS and Onboard Camera Data Describing Traffic Interactions

This project was part of the Characterizing Behaviors and Capabilities for Emerging Connected and Automated Vehicle Technologies, Sensors, and Connectivity project. The National Laboratory of the Rockies partnered with Cummins Inc. to collect data from Class 8 tractor trailer combinations in platoon (cooperative adaptive cruise control) operations on public roads in southern Indiana. Data collected include J1939 CAN bus, radar, intervehicle position, and video data. The video data could not be shared in the raw form, so they were processed to extract information on the other vehicles on the road, their relative positions, and intrusion events. This information was then columnized for modeling use and further enhanced by appending road information including road type, speed limit, altitude, and grade. The test route included free-flowing traffic, highway interchanges, and construction zones, as well as low-, medium-, and high-grade sections. Individual test conditions varied by day, with advanced driver-assistance system (ADAS) features engaged or disengaged and different combined vehicle masses tested in addition to uncontrolled variables such as weather and traffic interactions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

HAMscope: a snapshot Hyperspectral Autofluorescence Miniscope for real-time molecular imaging

We introduce HAMscope, a compact, snapshot hyperspectral autofluorescence miniscope that enables real-time, label-free molecular imaging in a wide range of biological systems. By integrating a thin polymer diffuser into a widefield miniscope, HAMscope spectrally encodes each frame and employs a probabilistic deep learning framework to reconstruct 30-channel hyperspectral stacks (452-703 nm) or directly infer molecular composition maps from single images. A scalable multi-pass U-Net architecture with transformer-based attention and per pixel uncertainty estimation enables high spatio-spectral fidelity (mean absolute error ∼0.0048) at video rates. While initially demonstrated in plant systems, including lignin, chlorophyll, and suberin imaging in intact poplar and cork tissues, the platform is readily adaptable to other applications such as neural activity mapping, metabolic profiling, and histopathology. We show that the system generalizes to out-of-distribution tissue types and supports direct molecular mapping without the need for spectral unmixing. HAMscope establishes a general framework for compact, uncertainty-aware spectral imaging that combines minimal optics with advanced deep learning, offering broad utility for real-time biochemical imaging across neuroscience, environmental monitoring, and biomedicine.

59 BASIC BIOLOGICAL SCIENCES↗

Measurement of Photovoltaic Module Deformation Dynamics During Hail Impact Using Digital Image Correlation

Stereo high-speed video of photovoltaic modules undergoing laboratory hail tests was processed using digital image correlation to determine module surface deformation during and immediately following impact. The purpose of this work was to demonstrate a methodology for characterizing module impact response differences as a function of construction and incident hail parameters. Video capture and digital image analysis were able to capture out-of-plane module deformation to a resolution of ±0.1 mm at 11 kHz on an in-plane grid of 10 × 10 mm over the area of a 1 × 2 m commercial photovoltaic module. With lighting and optical adjustments, the technique was adaptable to arbitrary module designs, including size, backsheet color, and cell interconnection. Furthermore, impacts were observed to produce an initially localized dimple in the glass surface, with peak deflection proportional to the square root of incident energy. Subsequent deformation propagation and dissipation were also captured, along with behavior for instances when the module glass fractured. Natural frequencies of the module were identifiable by analyzing module oscillations postimpact. Limitations of the measurement technique were that the impacting ice ball obscured the data field immediately surrounding the point of contact, and both ice and glass fracture events occurred within 100 μs, which was not resolvable at the chosen frame rate. Increasing the frame rate and visualizing the back surface of the impact could be applied to avoid these issues. Applications for these data include validating computational models for hail impacts, identifying the natural frequencies of a module, and identifying damage initiation mechanisms.

14 SOLAR ENERGY↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗

AMVOS: Additive Manufacturing Video Object Segmentation Dataset

This dataset provides labeled video frames from four additive manufacturing (AM) processes for video object segmentation (VOS) tasks. It contains 90 video segments comprising 900 individually annotated frames across five AM datasets: laser hot-wire directed energy deposition (LHW-DED), tungsten inert gas wire arc additive manufacturing (TIG-WAAM), plasma arc welding (PAW), visible-light polymer extrusion (visPolymer), and near-infrared polymer extrusion (irPolymer). Each video segment consists of 10 contiguous frames with corresponding pixel-level object instance annotations. Depending on the process, two of four object classes are labeled per frame: Melt Pool, Feed Wire, Nozzle, or Material. Raw frames are provided as .jpg files and annotations as palettized .png files. The dataset follows the directory structure of established VOS benchmarks (DAVIS, YouTube-VOS, MOSE), enabling direct integration into VOS model training and evaluation pipelines for foundation model fine-tuning, domain adaptation, or zero-shot performance benchmarking. Data was collected at Oak Ridge National Laboratory's Manufacturing Demonstration Facility.

Wetzel, Jon [ORNL]↗

Vehicle Localization in 3D World Coordinates Using Single Camera at Traffic Intersection

Optimizing traffic control systems at traffic intersections can reduce the network-wide fuel consumption, as well as emissions of conventional fuel-powered vehicles. While traffic signals have been controlled based on predetermined schedules, various adaptive signal control systems have recently been developed using advanced sensors such as cameras, radars, and LiDARs. Among these sensors, cameras can provide a cost-effective way to determine the number, location, type, and speed of the vehicles for better-informed decision-making at traffic intersections. In this research, a new approach for accurately determining vehicle locations near traffic intersections using a single camera is presented. For that purpose, a well-known object detection algorithm called YOLO is used to determine vehicle locations in video images captured by a traffic camera. YOLO draws a bounding box around each detected vehicle, and the vehicle location in the image coordinates is converted to the world coordinates using camera calibration data. During this process, a significant error between the center of a vehicle’s bounding box and the real center of the vehicle in the world coordinates is generated due to the angled view of the vehicles by a camera installed on a traffic light pole. As a means of mitigating this vehicle localization error, two different types of regression models are trained and applied to the centers of the bounding boxes of the camera-detected vehicles. The accuracy of the proposed approach is validated using both static camera images and live-streamed traffic video. Based on the improved vehicle localization, it is expected that more accurate traffic signal control can be made to improve the overall network-wide energy efficiency and traffic flow at traffic intersections.

47 OTHER INSTRUMENTATION↗

Bottleneck Detection in Modular Construction Factories Using Computer Vision

The construction industry is increasingly adopting off-site and modular construction methods due to the advantages offered in terms of safety, quality, and productivity for construction projects. Despite the advantages promised by this method of construction, modular construction factories still rely on manually-intensive work, which can lead to highly variable cycle times. As a result, these factories experience bottlenecks in production that can reduce productivity and cause delays to modular integrated construction projects. To remedy this effect, computer vision-based methods have been proposed to monitor the progress of work in modular construction factories. However, these methods fail to account for changes in the appearance of the modular units during production, they are difficult to adapt to other stations and factories, and they require a significant amount of annotation effort. Due to these drawbacks, this paper proposes a computer vision-based progress monitoring method that is easy to adapt to different stations and factories and relies only on two image annotations per station. In doing so, the Scale-invariant feature transform (SIFT) method is used to identify the presence of modular units at workstations, and the Mask R-CNN deep learning-based method is used to identify active workstations. This information was synthesized using a near real-time data-driven bottleneck identification method suited for assembly lines in modular construction factories. This framework was successfully validated using 420 h of surveillance videos of a production line in a modular construction factory in the U.S., providing 96% accuracy in identifying the occupancy of the workstations and an F-1 Score of 89% in identifying the state of each station on the production line. The extracted active and inactive durations were successfully used via a data-driven bottleneck detection method to detect bottleneck stations inside a modular construction factory. The implementation of this method in factories can lead to continuous and comprehensive monitoring of the production line and prevent delays by timely identification of bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION↗