<abstract xmlns="http://www.w3.org/1999/xhtml"><sec id="j_jdis-2019-0006_s_006_w2aab3b7b1b1b6b1aab1c17b1Aa"><h3>Purpose</h3><p>The ability to identify the scholarship of individual authors is essential for performance evaluation. A number of factors hinder this endeavor. Common and similarly spelled surnames make it difficult to isolate the scholarship of individual authors indexed on large databases. Variations in name spelling of individual scholars further complicates matters. Common family names in scientific powerhouses like China make it problematic to distinguish between authors possessing ubiquitous and/or anglicized surnames (as well as the same or similar first names). The assignment of unique author identifiers provides a major step toward resolving these difficulties. We maintain, however, that in and of themselves, author identifiers are not sufficient to fully address the author uncertainty problem. In this study we build on the author identifier approach by considering commonalities in fielded data between authors containing the same surname and first initial of their first name. We illustrate our approach using three case studies.</p></sec><sec id="j_jdis-2019-0006_s_007_w2aab3b7b1b1b6b1aab1c17b2Aa"><h3>Design/methodology/approach</h3><p>The approach we advance in this study is based on commonalities among fielded data in search results. We cast a broad initial net—i.e., a Web of Science (WOS) search for a given author’s last name, followed by a comma, followed by the first initial of his or her first name (e.g., a search for ‘John Doe’ would assume the form: ‘Doe, J’). Results for this search typically contain all of the scholarship legitimately belonging to this author in the given database (i.e., all of his or her true positives), along with a large amount of noise, or scholarship not belonging to this author (i.e., a large number of false positives). From this corpus we proceed to iteratively weed out false positives and retain true positives. Author identifiers provide a good starting point—e.g., if ‘Doe, J’ and ‘Doe, John’ share the same author identifier, this would be sufficient for us to conclude these are one and the same individual. We find email addresses similarly adequate—e.g., if two author names which share the same surname and same first initial have an email address in common, we conclude these authors are the same person. Author identifier and email address data is not always available, however. When this occurs, other fields are used to address the author uncertainty problem.</p><p>Commonalities among author data other than unique identifiers and email addresses is less conclusive for name consolidation purposes. For example, if ‘Doe, John’ and ‘Doe, J’ have an affiliation in common, do we conclude that these names belong the same person? They may or may not; affiliations have employed two or more faculty members sharing the same last and first initial. Similarly, it’s conceivable that two individuals with the same last name and first initial publish in the same journal, publish with the same co-authors, and/or cite the same references. Should we then ignore commonalities among these fields and conclude they’re too imprecise for name consolidation purposes? It is our position that such commonalities are indeed valuable for addressing the author uncertainty problem, but more so when used in combination.</p><p>Our approach makes use of automation as well as manual inspection, relying initially on author identifiers, then commonalities among fielded data other than author identifiers, and finally manual verification. To achieve name consolidation independent of author identifier matches, we have developed a procedure that is used with bibliometric software called VantagePoint (see <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.thevantagepoint.com">www.thevantagepoint.com)</ext-link> While the application of our technique does not exclusively depend on VantagePoint, it is the software we find most efficient in this study. The script we developed to implement this procedure is designed to implement our name disambiguation procedure in a way that significantly reduces manual effort on the user’s part. Those who seek to replicate our procedure independent of VantagePoint can do so by manually following the method we outline, but we note that the manual application of our procedure takes a significant amount of time and effort, especially when working with larger datasets.</p><p>Our script begins by prompting the user for a surname and a first initial (for any author of interest). It then prompts the user to select a WOS field on which to consolidate author names. After this the user is prompted to point to the name of the authors field, and finally asked to identify a specific author name (referred to by the script as the primary author) within this field whom the user knows to be a true positive (a suggested approach is to point to an author name associated with one of the records that has the author’s ORCID iD or email address attached to it).</p><p>The script proceeds to identify and combine all author names sharing the primary author’s surname and first initial of his or her first name who share commonalities in the WOS field on which the user was prompted to consolidate author names. This typically results in significant reduction in the initial dataset size. After the procedure completes the user is usually left with a much smaller (and more manageable) dataset to manually inspect (and/or apply additional name disambiguation techniques to).</p></sec><sec id="j_jdis-2019-0006_s_008_w2aab3b7b1b1b6b1aab1c17b3Aa"><h3>Research limitations</h3><p>Match field coverage can be an issue. When field coverage is paltry dataset reduction is not as significant, which results in more manual inspection on the user’s part. Our procedure doesn’t lend itself to scholars who have had a legal family name change (after marriage, for example). Moreover, the technique we advance is (sometimes, but not always) likely to have a difficult time dealing with scholars who have changed careers or fields dramatically, as well as scholars whose work is highly interdisciplinary.</p></sec><sec id="j_jdis-2019-0006_s_009_w2aab3b7b1b1b6b1aab1c17b4Aa"><h3>Practical implications</h3><p>The procedure we advance has the ability to save a significant amount of time and effort for individuals engaged in name disambiguation research, especially when the name under consideration is a more common family name. It is more effective when match field coverage is high and a number of match fields exist.</p></sec><sec id="j_jdis-2019-0006_s_010_w2aab3b7b1b1b6b1aab1c17b5Aa"><h3>Originality/value</h3><p>Once again, the procedure we advance has the ability to save a significant amount of time and effort for individuals engaged in name disambiguation research. It combines preexisting with more recent approaches, harnessing the benefits of both.</p></sec><sec id="j_jdis-2019-0006_s_011_w2aab3b7b1b1b6b1aab1c17b6Aa"><h3>Findings</h3><p>Our study applies the name disambiguation procedure we advance to three case studies. Ideal match fields are not the same for each of our case studies. We find that match field effectiveness is in large part a function of field coverage. Comparing original dataset size, the timeframe analyzed for each case study is not the same, nor are the subject areas in which they publish. Our procedure is more effective when applied to our third case study, both in terms of list reduction and 100% retention of true positives. We attribute this to excellent match field coverage, and especially in more specific match fields, as well as having a more modest/manageable number of publications.</p><p>While machine learning is considered authoritative by many, we do not see it as practical or replicable. The procedure advanced herein is both practical, replicable and relatively user friendly. It might be categorized into a space between ORCID and machine learning. Machine learning approaches typically look for commonalities among citation data, which is not always available, structured or easy to work with. The procedure we advance is intended to be applied across numerous fields in a dataset of interest (e.g. emails, coauthors, affiliations, etc.), resulting in multiple rounds of reduction. Results indicate that effective match fields include author identifiers, emails, source titles, co-authors and ISSNs. While the script we present is not likely to result in a dataset consisting solely of true positives (at least for more common surnames), it does significantly reduce manual effort on the user’s part. Dataset reduction (after our procedure is applied) is in large part a function of (a) field availability and (b) field coverage.</p></sec></abstract>

A Multi-match Approach to the Author Uncertainty Problem

<abstract xmlns="http://www.w3.org/1999/xhtml"><sec id="j_jdis-2019-0007_s_006_w2aab3b7b2b1b6b1aab1c17b1Aa"><h3>Purpose</h3><p>To design and test a method for normalizing book citations in Google Scholar.</p></sec><sec id="j_jdis-2019-0007_s_007_w2aab3b7b2b1b6b1aab1c17b2Aa"><h3>Design/methodology/approach</h3><p>A hybrid citing-side, cited-side normalization method was developed and this was tested on a sample of 285 research monographs. The results were analyzed and conclusions drawn.</p></sec><sec id="j_jdis-2019-0007_s_008_w2aab3b7b2b1b6b1aab1c17b3Aa"><h3>Findings</h3><p>The method was technically feasible but required extensive manual intervention because of the poor quality of the Google Scholar data.</p></sec><sec id="j_jdis-2019-0007_s_009_w2aab3b7b2b1b6b1aab1c17b4Aa"><h3>Research limitations</h3><p>The sample of books was limited and also all were from one discipline —business and management. Also, the method has only been tested on Google Scholar, it would be useful to test it on Web of Science or Scopus.</p></sec><sec id="j_jdis-2019-0007_s_010_w2aab3b7b2b1b6b1aab1c17b5Aa"><h3>Practical limitations</h3><p>Google Scholar is a poor source of data although it does cover a much wider range citation sources that other databases.</p></sec><sec id="j_jdis-2019-0007_s_011_w2aab3b7b2b1b6b1aab1c17b6Aa"><h3>Originality/value</h3><p>This is the first method that has been developed specifically for normalizing books which have so far not been able to be normalized.</p></sec></abstract>

Normalizing Book Citations in Google Scholar: A Hybrid Cited-side Citing-side Method

<abstract xmlns="http://www.w3.org/1999/xhtml"><sec id="j_jdis-2019-0008_s_006_w2aab3b7b3b1b6b1aab1c17b1Aa"><h3>Purpose</h3><p>The evolution of the socio-cognitive structure of the field of knowledge management (KM) during the period 1986–2015 is described.</p></sec><sec id="j_jdis-2019-0008_s_007_w2aab3b7b3b1b6b1aab1c17b2Aa"><h3>Design/methodology/approach</h3><p>Records retrieved from Web of Science were submitted to author co-citation analysis (ACA) following a longitudinal perspective as of the following time slices: 1986–1996, 1997–2006, and 2007–2015. The top 10% of most cited first authors by sub-periods were mapped in bibliometric networks in order to interpret the communities formed and their relationships.</p></sec><sec id="j_jdis-2019-0008_s_008_w2aab3b7b3b1b6b1aab1c17b3Aa"><h3>Findings</h3><p>KM is a homogeneous field as indicated by networks results. Nine classical authors are identified since they are highly co-cited in each sub-period, highlighting Ikujiro Nonaka as the most influential authors in the field. The most significant communities in KM are devoted to strategic management, KM foundations, organisational learning and behaviour, and organisational theories. Major trends in the evolution of the intellectual structure of KM evidence a technological influence in 1986–1996, a strategic influence in 1997–2006, and finally a sociological influence in 2007–2015.</p></sec><sec id="j_jdis-2019-0008_s_009_w2aab3b7b3b1b6b1aab1c17b4Aa"><h3>Research limitations</h3><p>Describing a field from a single database can offer biases in terms of output coverage. Likewise, the conference proceedings and books were not used and the analysis was only based on first authors. However, the results obtained can be very useful to understand the evolution of KM research.</p></sec><sec id="j_jdis-2019-0008_s_010_w2aab3b7b3b1b6b1aab1c17b5Aa"><h3>Practical implications</h3><p>These results might be useful for managers and academicians to understand the evolution of KM field and to (re)define research activities and organisational projects.</p></sec><sec id="j_jdis-2019-0008_s_011_w2aab3b7b3b1b6b1aab1c17b6Aa"><h3>Originality/value</h3><p>The novelty of this paper lies in considering ACA as a bibliometric technique to study KM research. In addition, our investigation has a wider time coverage than earlier articles.</p></sec></abstract>

Evolution of the Socio-cognitive Structure of Knowledge Management (1986–2015): An Author Co-citation Analysis

<abstract xmlns="http://www.w3.org/1999/xhtml"><sec id="j_jdis-2019-0009_s_006_w2aab3b7b4b1b6b1aab1c17b1Aa"><h3>Purpose</h3><p>Study how economic parameters affect positions in the Academic Ranking of World Universities’ top 500 published by the Shanghai Jiao Tong University Graduate School of Education in countries/regions with listed higher education institutions.</p></sec><sec id="j_jdis-2019-0009_s_007_w2aab3b7b4b1b6b1aab1c17b2Aa"><h3>Design/methodology/approach</h3><p>The methodology used capitalises on the multi-variate characteristics of the data analysed. The multi-colinearity problem posed is solved by running principal components prior to regression analysis, using both classical (OLS) and robust (Huber and Tukey) methods.</p></sec><sec id="j_jdis-2019-0009_s_008_w2aab3b7b4b1b6b1aab1c17b3Aa"><h3>Findings</h3><p>Our results revealed that countries/regions with long ranking traditions are highly competitive. Findings also showed that some countries/regions such as Germany, United Kingdom, Canada, and Italy, had a larger number of universities in the top positions than predicted by the regression model. In contrast, for Japan, a country where social and economic performance is high, the number of ARWU universities projected by the model was much larger than the actual figure. In much the same vein, countries/regions that invest heavily in education, such as Japan and Denmark, had lower than expected results.</p></sec><sec id="j_jdis-2019-0009_s_009_w2aab3b7b4b1b6b1aab1c17b4Aa"><h3>Research limitations</h3><p>Using data from only one ranking is a limitation of this study, but the methodology used could be useful to other global rankings.</p></sec><sec id="j_jdis-2019-0009_s_010_w2aab3b7b4b1b6b1aab1c17b5Aa"><h3>Practical implications</h3><p>The results provide good insights for policy makers. They indicate the existence of a relationship between research output and the number of universities per million inhabitants. Countries/regions, which have historically prioritised higher education, exhibited highest values for indicators that compose the rankings methodology; furthermore, minimum increase in welfare indicators could exhibited significant rises in the presence of their universities on the rankings.</p></sec><sec id="j_jdis-2019-0009_s_011_w2aab3b7b4b1b6b1aab1c17b6Aa"><h3>Originality/value</h3><p>This study is well defined and the result answers important questions about characteristics of countries/regions and their higher education system.</p></sec></abstract>

Does a Country/Region’s Economic Status Affect Its Universities’ Presence in International Rankings?

<abstract xmlns="http://www.w3.org/1999/xhtml"><sec id="j_jdis-2019-0010_s_005_w2aab3b7b5b1b6b1aab1c17b1Aa"><h3>Purpose</h3><p>To investigate the effectiveness of using node2vec on journal citation networks to represent journals as vectors for tasks such as clustering, science mapping, and journal diversity measure.</p></sec><sec id="j_jdis-2019-0010_s_006_w2aab3b7b5b1b6b1aab1c17b2Aa"><h3>Design/methodology/approach</h3><p>Node2vec is used in a journal citation network to generate journal vector representations.</p></sec><sec id="j_jdis-2019-0010_s_007_w2aab3b7b5b1b6b1aab1c17b3Aa"><h3>Findings</h3><p>1. Journals are clustered based on the node2vec trained vectors to form a science map. 2. The norm of the vector can be seen as an indicator of the diversity of journals. 3. Using node2vec trained journal vectors to determine the Rao-Stirling diversity measure leads to a better measure of diversity than that of direct citation vectors.</p></sec><sec id="j_jdis-2019-0010_s_008_w2aab3b7b5b1b6b1aab1c17b4Aa"><h3>Research limitations</h3><p>All analyses use citation data and only focus on the journal level.</p></sec><sec id="j_jdis-2019-0010_s_009_w2aab3b7b5b1b6b1aab1c17b5Aa"><h3>Practical implications</h3><p>Node2vec trained journal vectors embed rich information about journals, can be used to form a science map and may generate better values of journal diversity measures.</p></sec><sec id="j_jdis-2019-0010_s_010_w2aab3b7b5b1b6b1aab1c17b6Aa"><h3>Originality/value</h3><p>The effectiveness of node2vec in scientometric analysis is tested. Possible indicators for journal diversity measure are presented.</p></sec></abstract>

Node2vec Representation for Clustering Journals and as A Possible Measure of Diversity

AHEAD OF PRINT

Volume 9 (2024): Issue 2 (April 2024)

Volume 9 (2024): Issue 1 (February 2024)

Volume 8 (2023): Issue 4 (November 2023)

Volume 8 (2023): Issue 3 (June 2023)

Volume 8 (2023): Issue 2 (April 2023)

Volume 8 (2023): Issue 1 (February 2023)

Volume 7 (2022): Issue 4 (November 2022)

Volume 7 (2022): Issue 3 (August 2022)

Volume 7 (2022): Issue 2 (April 2022)

Volume 7 (2022): Issue 1 (February 2022)

Volume 6 (2021): Issue 4 (November 2021)

Volume 6 (2021): Issue 3 (June 2021)

Volume 6 (2021): Issue 2 (April 2021)

Volume 6 (2021): Issue 1 (February 2021)

Volume 5 (2020): Issue 4 (November 2020)

Volume 5 (2020): Issue 3 (August 2020)

Volume 5 (2020): Issue 2 (April 2020)

Volume 5 (2020): Issue 1 (February 2020)

Volume 4 (2019): Issue 4 (December 2019)

Volume 4 (2019): Issue 3 (August 2019)

Volume 4 (2019): Issue 2 (May 2019)

Volume 4 (2019): Issue 1 (February 2019)

Volume 3 (2018): Issue 4 (November 2018)

Volume 3 (2018): Issue 3 (August 2018)

Volume 3 (2018): Issue 2 (May 2018)

Volume 3 (2018): Issue 1 (February 2018)

Volume 2 (2017): Issue 4 (December 2017)

Volume 2 (2017): Issue 3 (August 2017)

Volume 2 (2017): Issue 2 (May 2017)

Volume 2 (2017): Issue 1 (February 2017)

Volume 1 (2016): Issue 4 (November 2016)

Volume 1 (2016): Issue 3 (August 2016)

Volume 1 (2016): Issue 2 (May 2016)

Volume 1 (2016): Issue 1 (February 2016)

Journal of Data and Information Science

Journal of Data and Information Science (JDIS, formerly Chinese Journal of Library and Information Science), sponsored by the Chinese Academy of Sciences (CAS) and published quarterly by the National Science Library of CAS, is the first internationally published English-language academic journal in Library and Information Science and related fields from China.  The Journal of Data and Information Science (JDIS) focuses on data-based research oriented toward the exploration of scientific research and innovation. The main areas of interest are science of science, evidence-based policymaking, research evaluation, computational social science, and scientometrics/bibliometrics/altmetrics/ informetrics. Emphasis is given to research that focuses on data, analytics, and knowledge discovery, and supports decision making and science policy. This includes modeling, innovation, data security, media and communications, and social development. Topics may include studies of metadata or full content data, text or non-textural data, structured or non-structural data, domain-specific or cross-domain data, and dynamic or interactive data.  Specific topic areas may include (but are not limited to):    Knowledge organization  Knowledge discovery and data mining  Knowledge integration and fusion  Semantic Web  Science of science  Bibliometrics and scientometrics  Analytic and diagnostic informetrics  Competitive intelligence  Predictive analysis  Social network analysis and metrics  Semantic and interactively analytic retrieval  Evidence-based policy analysis  Intelligent knowledge production  Knowledge-driven workflow management and decision-making  Knowledge-driven collaboration and its management  Domain knowledge infrastructure with knowledge fusion and analytics  Training for data &amp; information scientists  Development of data and information services    JDIS publishes theoretical and empirical work. Systematic reviews are welcome and applied research in development of advanced methods, services, and best practices is also an important part. But simple application of established informetrics on a specific research field or country is out of the scope.Welcome to submit your papers to JDIS.  Why subscribe and read  JDIS is the first and only English journal from China in Library and Information Science and related fields. With an aim to disseminate the cutting-edge research in these fields, it is devoted to the study and application of the theories, methods, techniques, services, and infrastructural facilities using big data to support knowledge discovery for decision and policy making. The basic emphasis is big data-based, analytics centered, knowledge discovery driven, and decision making supporting. JDIS has gathered a big body of high profile experts across the world who contribute their research to the journal. The international authors account for around 62% in its first publication year (2016).  Why submit  JDIS is the first and only English journal from China in Library and Information Science and related fields. It owns a number of world front-line scholars as editorial board members or reviewers. The turnaround time on average for a manuscript from submission to final decision is less than two and a half months.  Archiving  Sciendo archives the contents of this journal in Portico- digital long-term preservation service of scholarly books, journals and collections.  Plagiarism Policy  The editorial board is participating in a growing community of Similarity Check System's users in order to ensure that the content published is original and trustworthy. Similarity Check is a medium that allows for comprehensive manuscripts screening, aimed to eliminate plagiarism and provide a high standard and quality peer-review process.