Anmeldung Registrierung
Auto Hell Dunkel
Erweiterte Suche
  1. Startseite
  2. Podcasts
  3. Medizin - Open Access LMU - Teil 15/22 Podcast
  4. Bias in random forest variable importance measures: Illustrations, sources and a solution
Background: Variable importance measures for random forests have
been receiving increased attention as a means of variable selection
in many classification tasks in bioinformatics and related
scientific fields, for instance to select a subset of genetic
markers relevant for the prediction of a certain disease. We show
that random forest variable importance measures are a sensible
means for variable selection in many applications, but are not
reliable in situations where potential predictor variables vary in
their scale of measurement or their number of categories. This is
particularly important in genomics and computational biology, where
predictors often include variables of different types, for example
when predictors include both sequence data and continuous variables
such as folding energy, or when amino acid sequence data show
different numbers of categories. Results: Simulation studies are
presented illustrating that, when random forest variable importance
measures are used with data of varying types, the results are
misleading because suboptimal predictor variables may be
artificially preferred in variable selection. The two mechanisms
underlying this deficiency are biased variable selection in the
individual classification trees used to build the random forest on
one hand, and effects induced by bootstrap sampling with
replacement on the other hand. Conclusion: We propose to employ an
alternative implementation of random forests, that provides
unbiased variable selection in the individual classification trees.
When this method is applied using subsampling without replacement,
the resulting variable importance measures can be used reliably for
variable selection even in situations where the potential predictor
variables vary in their scale of measurement or their number of
categories. The usage of both random forest algorithms and their
variable importance measures in the R system for statistical
computing is illustrated and documented thoroughly in an
application re-analyzing data from a study on RNA editing.
Therefore the suggested method can be applied straightforwardly by
scientists in bioinformatics research.
Episode melden

„Bias in random forest variable importance measures: Illustrations, sources and a solution“

Worum geht es? Danach fragen wir noch nach dem Grund.

Abonnenten

Teilen

Mein Archiv

Deine Privatkopie der Folgen, die du nicht verlieren willst.

Podcast-Folgen verschwinden. Feeds werden auf die letzten Episoden gekürzt, Hoster räumen alte Dateien ab, Formate wechseln den Anbieter und lassen ihr Archiv zurück. Mit „Mein Archiv“ sichert podcast.de die Folgen deiner Podcasts für dich — angefangen bei den ältesten, denn die sind zuerst weg.

  • Deine gesicherten Folgen bleiben hörbar, auch wenn das Original offline geht.
  • Auch Folgen, die im heutigen Feed gar nicht mehr stehen — podcast.de kennt sie noch.
  • Herunterladen bleibt möglich, solange die Folge beim Podcaster liegt. Der zählt seine Abrufe wie bisher.
Startet bald

Sei beim Start von Mein Archiv dabei

Mein Archiv ist fast fertig. Trag dich ein, dann bekommst du eine E-Mail, sobald es losgeht – und bist von Anfang an dabei. Wir schreiben dir nur zum Start, keine Werbung, keine Weitergabe deiner Daten.

Du bekommst zuerst eine Bestätigungsmail. Abmelden geht jederzeit. Datenschutz