Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Learning from Crowds by Modeling Common Confusions

Crowdsourcing provides a practical way to obtain large amounts of labeled data at a low cost. However, the annotation quality of annotators varies considerably, which imposes new challenges in learning a high-quality model from the crowdsourced annotations. In this work, we provide a new perspective to decompose annotation noise into common noise and individual noise and differentiate the source of confusion based on instance difficulty and annotator expertise on a per-instance-annotator basis. We realize this new crowdsourcing model by an end-to-end learning solution with two types of noise adaptation layers: one is shared across annotators to capture their commonly shared confusions, and the other one is pertaining to each annotator to realize individual confusion. To recognize the source of noise in each annotation, we use an auxiliary network to choose from the two noise adaptation layers with respect to both instances and annotators. Extensive experiments on both synthesized and real-world benchmarks demonstrate the effectiveness of our proposed common noise adaptation solution.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Generating Customized Verifiers for Automatically Generated Code

Program verification using Hoare-style techniques requires many logical annotations. We have previously developed a generic annotation inference algorithm that weaves in all annotations required to certify safety properties for automatically generated code. It uses patterns to capture generator- and property-specific code idioms and property-specific meta-program fragments to construct the annotations. The algorithm is customized by specifying the code patterns and integrating them with the meta-program fragments for annotation construction. However, this is difficult since it involves tedious and error-prone low-level term manipulations. Here, we describe an annotation schema compiler that largely automates this customization task using generative techniques. It takes a collection of high-level declarative annotation schemas tailored towards a specific code generator and safety property, and generates all customized analysis functions and glue code required for interfacing with the generic algorithm core, thus effectively creating a customized annotation inference algorithm. The compiler raises the level of abstraction and simplifies schema development and maintenance. It also takes care of some more routine aspects of formulating patterns and schemas, in particular handling of irrelevant program fragments and irrelevant variance in the program structure, which reduces the size, complexity, and number of different patterns and annotation schemas that are required. The improvements described here make it easier and faster to customize the system to a new safety property or a new generator, and we demonstrate this by customizing it to certify frame safety of space flight navigation code that was automatically generated from Simulink models by MathWorks' Real-Time Workshop.

Denney, Ewen↗

The Gene Ontology resource: enriching a GOld mine

The Gene Ontology Consortium (GOC) provides the most comprehensive resource currently available for computable knowledge regarding the functions of genes and gene products. Here, we report the advances of the consortium over the past two years. The new GO-CAM annotation framework was notably improved, and we formalized the model with a computational schema to check and validate the rapidly increasing repository of 2838 GO-CAMs. In addition, we describe the impacts of several collaborations to refine GO and report a 10% increase in the number of GO annotations, a 25% increase in annotated gene products, and over 9,400 new scientific articles annotated. As the project matures, we continue our efforts to review older annotations in light of newer findings, and, to maintain consistency with other ontologies. As a result, 20,000 annotations derived from experimental data were reviewed, corresponding to 2.5% of experimental GO annotations. The website (http://geneontology.org) was redesigned for quick access to documentation, downloads and tools. To maintain an accurate resource and support traceability and reproducibility, we have made available a historical archive covering the past 15 years of GO data with a consistent format and file structure for both the ontology and annotations.

59 BASIC BIOLOGICAL SCIENCES↗

DRAM example narrative

DRAM example narrative DRAM on KBase let's anyone run annotations using DRAM in the cloud. DRAM is an annotation tool that can annotate bacterial, archaeal and viral genomes and distills those annotatios into represetations of the functional genomic potential of those organisms. If you want to read more about DRAM you can check out the GitHub, wiki and journal article. DRAM annotate assemblies In KBase Assembly objects contain nucleotide sequences from genomes or metagenomes. DRAM can predict genes and annotate their function from KBase Assembly objects which may be microbial isolate genomes, metagenome assembled genomes or metagenomes. This is done with the Annotate and Distill Assemblies with DRAM app. This app can also anntoate AssemblySet objects which contain collection of Assembly objects. It also generates a Genome object and a GenomeSet object which can be used for further analysis with other KBase apps. The full annotations and other DRAM files are also available for download in the app.

59 BASIC BIOLOGICAL SCIENCES↗

Are Phosphatidic Acids Ubiquitous in Mammalian Tissues or Overemphasized in Mass Spectrometry Imaging Applications?

Abstract Mass spectrometry imaging (MSI) is an invaluable tool for the spatial visualization of molecules in vivo. However, the question of whether observed annotations are endogenous or artificial (i. e., from in‐source fragmentation) is critical and has been largely unexplored in multimodal MSI. In matrix‐assisted laser desorption/ionization (MALDI)‐MSI datasets from researchers worldwide, PAs were found to represent up to 18 % of annotations in rat brain. Rat brain was additionally imaged here using nanospray desorption electrospray ionization (nano‐DESI), a softer ionization strategy. No PAs observed with MALDI were present in the nano‐DESI dataset. Further investigation strongly indicated lipid fragmentation to PAs for MALDI‐MSI, but not with nano‐DESI‐MSI. We finally extend this observation to the MALDI‐MSI analyses of human tissues, showing that PA annotations comprised up to 16 % of annotations. Therefore, this study shows that MSI annotations should be carefully interrogated, as in‐source fragmentation or modification of lipids may contribute substantially to false annotations and incorrect biological interpretations.

Vandergrift, Gregory W.↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Improve Learning from Crowds via Generative Augmentation

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotators. This creates a sparsity issue and limits the quality of machine learning models trained on such data. In this paper, we study how to handle sparsity in crowdsourced data using data augmentation. Specifically, we propose to directly learn a classifier by augmenting the raw sparse annotations. We implement two principles of high-quality augmentation using Generative Adversarial Networks: 1) the generated annotations should follow the distribution of authentic ones, which is measured by a discriminator; 2) the generated annotations should have high mutual information with the ground-truth labels, which is measured by an auxiliary network. Extensive experiments and comparisons against an array of state-of-the-art learning from crowds methods on three real-world datasets proved the effectiveness of our data augmentation framework. It shows the potential of our algorithm for low-budget crowdsourcing in general.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Improvement of eukaryotic protein predictions from soil metagenomes

During the last decades, metagenomics has highlighted the diversity of microorganisms from environmental or host-associated samples. Most metagenomics public repositories use annotation pipelines tailored for prokaryotes regardless of the taxonomic origin of contigs. Consequently, eukaryotic contigs with intrinsically different gene features, are not optimally annotated. Using a bioinformatics pipeline, we have filtered 7.9 billion contigs from 6,872 soil metagenomes in the JGI’s IMG/M database to identify eukaryotic contigs. We have re-annotated genes using eukaryote-tailored methods, yielding 8 million eukaryotic proteins and over 300,000 orphan proteins lacking homology in public databases. Comparing the gene predictions we made with initial JGI ones on the same contigs, we confirmed our pipeline improves eukaryotic proteins completeness and contiguity in soil metagenomes. The improved quality of eukaryotic proteins combined with a more comprehensive assignment method yielded more reliable taxonomic annotation. This dataset of eukaryotic soil proteins with improved completeness, quality and taxonomic annotation reliability is of interest for any scientist aiming at studying the composition, biological functions and gene flux in soil communities involving eukaryotes.

54 ENVIRONMENTAL SCIENCES↗

EXSCLAIM!: Harnessing materials science literature for self-labeled microscopy datasets

This work introduces the EXSCLAIM! toolkit for the automatic extraction, separation, and caption-based natural language annotation of images from scientific literature. EXSCLAIM! is used to show how rule-based natural language processing and image recognition can be leveraged to construct an electron microscopy data set containing thousands of keyword-annotated nanostructure images. Moreover, it is demonstrated how a combination of statistical topic modeling and semantic word similarity comparisons can be used to increase the number and variety of keyword annotations on top of the standard annotations from EXSCLAIM! With large-scale imaging datasets constructed from scientific literature, users are well positioned to train neural networks for classification and recognition tasks specific to microscopy-tasks often otherwise inhibited by a lack of sufficient annotated training data.

36 MATERIALS SCIENCE↗

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

KBase Narrative - Porphyromonadaceae sp. W3.11 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗

KBase Narrative - Lachnospiraceae sp. C1.1 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗

Synthesizing Certified Code

Code certification is a lightweight approach to demonstrate software quality on a formal level. Its basic idea is to require producers to provide formal proofs that their code satisfies certain quality properties. These proofs serve as certificates which can be checked independently. Since code certification uses the same underlying technology as program verification, it also requires many detailed annotations (e.g., loop invariants) to make the proofs possible. However, manually adding theses annotations to the code is time-consuming and error-prone. We address this problem by combining code certification with automatic program synthesis. We propose an approach to generate simultaneously, from a high-level specification, code and all annotations required to certify generated code. Here, we describe a certification extension of AUTOBAYES, a synthesis tool which automatically generates complex data analysis programs from compact specifications. AUTOBAYES contains sufficient high-level domain knowledge to generate detailed annotations. This allows us to use a general-purpose verification condition generator to produce a set of proof obligations in first-order logic. The obligations are then discharged using the automated theorem E-SETHEO. We demonstrate our approach by certifying operator safety for a generated iterative data classification program without manual annotation of the code.

Whalen, Michael↗