Anmeldung Registrierung
Auto Hell Dunkel
Erweiterte Suche
  1. Startseite
  2. Podcasts
  3. Medizin - Open Access LMU - Teil 14/22 Podcast
  4. Gene and protein nomenclature in public databases
Background: Frequently, several alternative names are in use for
biological objects such as genes and proteins. Applications like
manual literature search, automated text-mining, named entity
identification, gene/protein annotation, and linking of knowledge
from different information sources require the knowledge of all
used names referring to a given gene or protein. Various
organism-specific or general public databases aim at organizing
knowledge about genes and proteins. These databases can be used for
deriving gene and protein name dictionaries. So far, little is
known about the differences between databases in terms of size,
ambiguities and overlap. Results: We compiled five gene and protein
name dictionaries for each of the five model organisms ( yeast,
fly, mouse, rat, and human) from different organism-specific and
general public databases. We analyzed the degree of ambiguity of
gene and protein names within and between dictionaries, to a
lexicon of common English words and domain-related non-gene terms,
and we compared different data sources in terms of size of
extracted dictionaries and overlap of synonyms between those. The
study shows that the number of genes/proteins and synonyms covered
in individual databases varies significantly for a given organism,
and that the degree of ambiguity of synonyms varies significantly
between different organisms. Furthermore, it shows that, despite
considerable efforts of co-curation, the overlap of synonyms in
different data sources is rather moderate and that the degree of
ambiguity of gene names with common English words and
domain-related non-gene terms varies depending on the considered
organism. Conclusion: In conclusion, these results indicate that
the combination of data contained in different databases allows the
generation of gene and protein name dictionaries that contain
significantly more used names than dictionaries obtained from
individual data sources. Furthermore, curation of combined
dictionaries considerably increases size and decreases ambiguity.
The entries of the curated synonym dictionary are available for
manual querying, editing, and PubMed- or Google-search via the
ProThesaurus-wiki. For automated querying via custom software, we
offer a web service and an exemplary client application.
Episode melden

„Gene and protein nomenclature in public databases“

Worum geht es? Danach fragen wir noch nach dem Grund.

Abonnenten

Teilen

Mein Archiv

Deine Privatkopie der Folgen, die du nicht verlieren willst.

Podcast-Folgen verschwinden. Feeds werden auf die letzten Episoden gekürzt, Hoster räumen alte Dateien ab, Formate wechseln den Anbieter und lassen ihr Archiv zurück. Mit „Mein Archiv“ sichert podcast.de die Folgen deiner Podcasts für dich — angefangen bei den ältesten, denn die sind zuerst weg.

  • Deine gesicherten Folgen bleiben hörbar, auch wenn das Original offline geht.
  • Auch Folgen, die im heutigen Feed gar nicht mehr stehen — podcast.de kennt sie noch.
  • Herunterladen bleibt möglich, solange die Folge beim Podcaster liegt. Der zählt seine Abrufe wie bisher.
Startet bald

Sei beim Start von Mein Archiv dabei

Mein Archiv ist fast fertig. Trag dich ein, dann bekommst du eine E-Mail, sobald es losgeht – und bist von Anfang an dabei. Wir schreiben dir nur zum Start, keine Werbung, keine Weitergabe deiner Daten.

Du bekommst zuerst eine Bestätigungsmail. Abmelden geht jederzeit. Datenschutz