Engineering PapersSearch

DOE OSTI · 2587970

Data and Code for Understanding Generative AI Content with Embedding Models

Abstract

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vargas, Max [Pacific Northwest National Laboratory (PNNL), Richland, WA (United States)], Engel, Andrew W [Pacific Northwest National Laboratory (PNNL), Richland, WA (United States)] (ORCID:000000032348483X), Chiang, Tony Y [Pacific Northwest National Laboratory (PNNL), Richland, WA (United States)], Cannon, Reilly T [Pacific Northwest National Laboratory (PNNL), Richland, WA (United States)], Sarwate, Anand D [Rutgers Univ., Piscataway, NJ (United States)]. 2025-08-25. Data and Code for Understanding Generative AI Content with Embedding Models. https://doi.org/10.25584/2587970

Cite the original work for its findings. Save a collection to share your selection of sources.