Characteristics of scientific Web publications: Preliminary data gathering and analysis

Because of the increasing presence of scientific publications on the Web, combined with the existing difficulties in easily verifying and retrieving these publications, research on techniques and methods for retrieval of scientific Web publications is called for. In this article, we report on the in...

Full description

Saved in:

Bibliographic Details
Published in	Journal of the American Society for Information Science and Technology Vol. 55; no. 14; pp. 1239 - 1249
Main Authors	Jepsen, Erik Thorlund, Seiden, Piet, Ingwersen, Peter, Björneborn, Lennart, Borlund, Pia
Format	Journal Article
Language	English
Published	Hoboken Wiley Subscription Services, Inc., A Wiley Company 01.12.2004 Wiley Periodicals Inc
Subjects	Accessibility Algorithms Bibliometrics Biology Botany Citation analysis Content analysis Content management systems Cutoffs Data analysis Documents Electronic publishing Feasibility studies Information retrieval Information Science Education Information Seeking Internet Libraries Library and information science Metadata Periodicals Plant biology Scholarly publishing Schools of library and information science Science Science and technology Scientific papers Search engines Search Strategies Studies URLs Webometrics Websites World Wide Web Denmark
Online Access	Get full text
ISSN	1532-2882 2330-1635 1532-2890 2330-1643
DOI	10.1002/asi.20079

Cover

Loading…

More Information
Summary:	Because of the increasing presence of scientific publications on the Web, combined with the existing difficulties in easily verifying and retrieving these publications, research on techniques and methods for retrieval of scientific Web publications is called for. In this article, we report on the initial steps taken toward the construction of a test collection of scientific Web publications within the subject domain of plant biology. The steps reported are those of data gathering and data analysis aiming at identifying characteristics of scientific Web publications. The data used in this article were generated based on specifically selected domain topics that are searched for in three publicly accessible search engines (Google, AllTheWeb, and AltaVista). A sample of the retrieved hits was analyzed with regard to how various publication attributes correlated with the scientific quality of the content and whether this information could be employed to harvest, filter, and rank Web publications. The attributes analyzed were inlinks, outlinks, bibliographic references, file format, language, search engine overlap, structural position (according to site structure), and the occurrence of various types of metadata. As could be expected, the ranked output differs between the three search engines. Apparently, this is caused by differences in ranking algorithms rather than the databases themselves. In fact, because scientific Web content in this subject domain receives few inlinks, both AltaVista and AllTheWeb retrieved a higher degree of accessible scientific content than Google. Because of the search engine cutoffs of accessible URLs, the feasibility of using search engine output for Web content analysis is also discussed.
Bibliography:	ark:/67375/WNG-VJ4FDD52-P ArticleID:ASI20079 istex:519DD96E20579E411A29541961202C899C1F67F6 ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14 ObjectType-Article-2 ObjectType-Feature-1 content type line 23
ISSN:	1532-2882 2330-1635 1532-2890 2330-1643
DOI:	10.1002/asi.20079