In Wolfgang Glänzel, Henk F. Moed, Ulrich Schmoch & Mike Thelwall (eds.), Springer Handbook of Science and Technology Indicators. Springer Verlag. pp. 781-800 (2019)
AbstractThis chapter reviews the development of data collection procedures on the web with an emphasis on current practices, data cleansing and matching, data quality and transparency. There are several issues to be considered when collecting data from the web. Transparency is essential to know what is included in the data source, how recent and comprehensive the data are, what timeframe is covered etc. Data quality relates to reliability and accuracy. Mistakes are inevitable, data providers, aggregators, and researchers all make mistakes, but these mistakes should be reduced to a minimum so that meaningful conclusions may be reached from the data analysis. Extensive data cleansing before starting the analysis is needed to try to correct mistakes in the data. When several data sources are used, data from different sources should be matched, and duplicates should be removed.
Similar books and articles
Using ‘new’ data sources for ‘old’ newspaper research: Developing guidelines for data collection.Peer Scheepers, Fred Wester & Pytrik Schafraad - 2006 - Communications 31 (4):455-467.
Internet-Based Data Collection: Promises and Realities.Jacob A. Benfield & William J. Szlemko - 2006 - Journal of Research Practice 2 (2):Article D1.
The Importance of Defining ‘Data’ in Data Management Policies: Commentary on: “Issues in Data Management”.Julie Richardson & Diane Hoffman-Kim - 2010 - Science and Engineering Ethics 16 (4):749-751.
The Time of Data: Timescales of Data Use in the Life Sciences.Sabina Leonelli - 2018 - Philosophy of Science 85 (5):741-754.
Data models and the acquisition and manipulation of data.Todd Harris - 2003 - Philosophy of Science 70 (5):1508-1517.
Data collection, counterterrorism and the right to privacy.Isaac Taylor - 2017 - Politics, Philosophy and Economics 16 (3):326-346.
Openness in Big Data and Data Repositories: The Application of an Ethics Framework for Big Data in Health and Research.Vicki Xafis & Markus K. Labude - 2019 - Asian Bioethics Review 11 (3):255-273.
Good Data.Angela Daly, Monique Mann & S. Kate Devitt - 2019 - Amsterdam, Netherlands: Institute of Network Cultures.
Data as oil, infrastructure or asset? Three metaphors of data as economic value.Jan Michael Nolin - 2019 - Journal of Information, Communication and Ethics in Society 18 (1):28-43.
Using models to correct data: paleodiversity and the fossil record.Alisa Bokulich - 2018 - Synthese 198 (Suppl 24):5919-5940.
The ethics of uncertainty for data subjects.Philip Nickel - 2019 - In Jenny Krutzinna & Luciano Floridi (eds.), The Ethics of Medical Data Donation. Springer Verlag. pp. 55-74.
Genomic Data-Sharing Practices.Angela G. Villanueva, Robert Cook-Deegan, Jill O. Robinson, Amy L. McGuire & Mary A. Majumder - 2019 - Journal of Law, Medicine and Ethics 47 (1):31-40.
Ethical Issues and Guidelines for Conducting Data Analysis in Psychological Research.Rachel Wasserman - 2013 - Ethics and Behavior 23 (1):3-15.
Added to PP
Historical graph of downloads