Information, learning and falsification

David Balduzzi

Information, learning and falsification

Abstract

There are (at least) three approaches to quantifying information. The first, algorithmic information or Kolmogorov complexity, takes events as strings and, given a universal Turing machine, quantifies the information content of a string as the length of the shortest program producing it [1]. The second, Shannon information, takes events as belonging to ensembles and quantifies the information resulting from observing the given event in terms of the number of alternate events that have been ruled out [2]. The third, statistical learning theory, has introduced measures of capacity that control (in part) the expected risk of classifiers [3]. These capacities quantify the expectations regarding future data that learning algorithms embed into classifiers. Solomonoff and Hutter have applied algorithmic information to prove remarkable results on universal induction. Shannon information provides the mathematical foundation for communication and coding theory. However, both approaches have shortcomings. Algorithmic information is not computable, severely limiting its practical usefulness. Shannon information refers to ensembles rather than actual events: it makes no sense to compute the Shannon information of a single string – or rather, there are many answers to this question depending on how a related ensemble is constructed. Although there are asymptotic results linking algorithmic and Shannon information, it is unsatisfying that there is such a large gap – a difference in kind – between the two measures. This note describes a new method of quantifying information, effective information, that links algorithmic information to Shannon information, and also links both to capacities arising in statistical learning theory [4, 5]. After introducing the measure, we show that it provides a non-universal analog of Kolmogorov complexity. We then apply it to derive basic capacities in statistical learning theory: empirical VC-entropy and empirical Rademacher complexity. A nice byproduct of our approach is an interpretation of the explanatory power of a learning algorithm in terms of the number of hypotheses it falsifies [6], counted in two different ways for the two capacities. We also discuss how effective information relates to information gain, Shannon and mutual information.

Cite

Plain text

BibTeX

Formatted text

Zotero

EndNote

Reference Manager

RefWorks

Options

Mark as duplicate

Find it on Scholar

Request removal from index

Revision history

Edit

Author's Profile

David Balduzzi

University of Zürich

Keywords

falsification information theory statistical learning theory kolmogorov complexity induction falsifiability Popper

Reprint years

My notes

Analytics

Added to PP
2013-09-26

Downloads
300 (#64,160)

6 months
46 (#84,461)

Historical graph of downloads

How can I increase my downloads?

Author's Profile

David Balduzzi

University of Zürich

Citations of this work

Falsifiable implies Learnable.David Balduzzi - manuscript

Add more citations

References found in this work

The Logic of Scientific Discovery.Karl Popper - 1959 - Studia Logica 9:262-265.

The Logic of Scientific Discovery.K. Popper - 1959 - British Journal for the Philosophy of Science 10 (37):55-57.

Causality: Models, Reasoning and Inference.Judea Pearl - 2000 - Tijdschrift Voor Filosofie 64 (1):201-202.

The Logic of Scientific Discovery.Karl R. Popper - 1959 - Les Etudes Philosophiques 14 (3):383-383.

Causality: Models, Reasoning and Inference.Christopher Hitchcock & Judea Pearl - 2001 - Philosophical Review 110 (4):639.

View all 7 references / Add more references

Applied ethics	Epistemology	History of Western Philosophy	Meta-ethics	Metaphysics	Normative ethics
Philosophy of biology	Philosophy of language	Philosophy of mind	Philosophy of religion	Science Logic and Mathematics	More ...

Information, learning and falsification

Abstract

Author's Profile

Categories

Keywords

Reprint years

Links

PhilArchive

External links

Through your library

My notes

Similar books and articles

Analytics

Author's Profile

Citations of this work

References found in this work