Engineering Papers⌕ Search

Engineering topics

Henry, Michael J.

Publications and source records attributed to Henry, Michael J..

Evaluating generative networks using Gaussian mixtures of image features

We develop a measure for evaluating the performance of generative networks given two sets of images. A popular performance measure currently used to do this is the Fréchet Inception Distance (FID). However, FID assumes that images featurized using the penultimate layer of Inception follow a Gaussian distribution. This assumption allows FID to be easily computed, since FID uses the 2-Wasserstein distance of two Gaussian distributions fitted to the featurized images. However, we show that Inception features of the ImageNet dataset are not Gaussian; in particular, each marginal is not Gaussian. To remedy this problem, we model the featurized images using Gaussian mixture models (GMMs) and compute the 2-Wasserstein distance restricted to GMMs. We define a performance measure, which we call WaM, on two sets of images by using inception (or another classifier) to featurize the images, estimate two GMMs, and use the restricted 2-Wasserstein distance to compare the GMMs. We experimentally show the advantages of WaM over FID, including how FID is more sensitive than WaM to image perturbations. By modelling the non-Gaussian features obtained from inception as GMMs and using a GMM metric, we can more accurately evaluate generative network performance.

machine learning, genrative adversarial networks↗

Video Summarization Using Deep Action Recognition Features and Robust Principal Components Analysis

In an instance where desired pre-defined actions, behaviors, or other categories are known a priori, various video classification and recognition models can be trained to discover those classifications and their location within the video. Absent that information, one might still be tasked with identifying interesting portions within a video, a process which—if done manually—is onerous and time-consuming as it requires manual inspection of the video itself. Recognizing high-level interesting segments within a whole video has been a general area of interest due to the ubiquity of video data. However the size of the data makes storage, retrieval, and inspection of large collections of videos cumbersome. This problem motivates the task of generating shortened clips highlighting the primary content of a video, relieving the burden of having to watch the entire video. This paper presents an unsupervised method of creating shortened clips of videos, enabling the rapid review of the most interesting content within a video. Our method uses features extracted from pre-trained action recognition models as input to online moving window robust principal component analysis to generate summaries. The procedure is tested on a publicly available video summarization dataset and demonstrates comparable performance to state-of-the-art in an un-augmented setting while requiring no training.

Claborne, Daniel M.↗

Probing for Artifacts: Detecting Imagenet Model Evasions

While deep learning models have made incredible progress across a variety of machine learning tasks, they remain vulnerable to adversarial examples crafted to fool otherwise trustworthy models. In this work we approach this problem through the lens of a detection framework. We propose a classification network that uses the hidden layer activations of a trained model as inputs to detect adversarial artifacts in an input. We train this classification network simultaneously against multiple adversarial algorithms to create a more robust detector and show higher detection rates than several alternatives. The novelty of our approach is in the scale and scope of probing Imagenet models for adversarial artifacts. In addition, we propose an improvement to feature squeezing, another common adversarial example detection method.

Rounds, Jeremiah↗