Dear all,
With the aim to compare Wikipedia traffic report data (e.g. viewing versus editing, regional differences within a language version, etc.), I have made a few more interactive infographics which show the historical changes since late 2011. (Historical numbers are scraped from the past versions archived by the Internet archive)
For more, please visit follow the link below: http://people.oii.ox.ac.uk/hanteng/2014/05/16/wikipedia-traffic/
It has at least one nice interactive feature: a user can zoom and pan to view the chart easily with a mouse or mousepad. The SVG vector-based presentation insures the picture quality is consistent when users zoom in to compare data points. (I haven't figured out how mpld3's html tooltip work for this project, though.)
It is also possible to extend the prototype with dynamic json objects so that the chart/tables can be updated automatically.
Any suggestions and comments are welcome.
Best, han-teng liao
Very useful Han-Teng, but one should note that the original data is about "the percentage of requesting ip addresses", excluding duplications of a single IP address within the same day, and not for example the number of edits. These two can be very different depending on dynamic/static IP address models in different countries. And that explains the discrepancy between your results and our earlier analysis based on circadian patterns and edits timestamps. http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0030091#pone-0030091-g004
Again, very interesting and well done. Best, Taha
On Fri, May 16, 2014 at 4:28 PM, h hanteng@gmail.com wrote:
Dear all,
With the aim to compare Wikipedia traffic report data (e.g. viewing versus editing, regional differences within a language version, etc.), I have made a few more interactive infographics which show the historical changes since late 2011. (Historical numbers are scraped from the past versions archived by the Internet archive)
For more, please visit follow the link below: http://people.oii.ox.ac.uk/hanteng/2014/05/16/wikipedia-traffic/
It has at least one nice interactive feature: a user can zoom and pan to view the chart easily with a mouse or mousepad. The SVG vector-based presentation insures the picture quality is consistent when users zoom in to compare data points. (I haven't figured out how mpld3's html tooltip work for this project, though.)
It is also possible to extend the prototype with dynamic json objectsso that the chart/tables can be updated automatically.
Any suggestions and comments are welcome.Best, han-teng liao
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
+ analytics
On May 16, 2014, at 8:50 AM, Taha Yasseri taha.yaseri@gmail.com wrote:
Very useful Han-Teng, but one should note that the original data is about "the percentage of requesting ip addresses", excluding duplications of a single IP address within the same day, and not for example the number of edits. These two can be very different depending on dynamic/static IP address models in different countries. And that explains the discrepancy between your results and our earlier analysis based on circadian patterns and edits timestamps.
Again, very interesting and well done. Best, Taha
On Fri, May 16, 2014 at 4:28 PM, h hanteng@gmail.com wrote: Dear all,
With the aim to compare Wikipedia traffic report data (e.g. viewing versus editing, regional differences within a language version, etc.), I have made a few more interactive infographics which show the historical changes since late 2011. (Historical numbers are scraped from the past versions archived by the Internet archive)
For more, please visit follow the link below: http://people.oii.ox.ac.uk/hanteng/2014/05/16/wikipedia-traffic/
It has at least one nice interactive feature: a user can zoom and pan to view the chart easily with a mouse or mousepad. The SVG vector-based presentation insures the picture quality is consistent when users zoom in to compare data points. (I haven't figured out how mpld3's html tooltip work for this project, though.)
It is also possible to extend the prototype with dynamic json objects so that the chart/tables can be updated automatically. Any suggestions and comments are welcome.Best, han-teng liao
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
-- .t _______________________________________________ Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
Thanks Taha for pointing that out. I have added the note to the blog post and am hope to start a conversation on what can we do to make the analysis and curation of Wikipedia traffic data more useful and meaningful both for research and policies.
BTW, very interesting phenomena of "sleep depth" for different languages in Taha Yasseri, Robert Sumi, János Kertész's paper. It provides insights into the time distribution of Wikipedia labour across working hours and working days. To certain extent, it shows us the current utility of the global "cognitive surplus" by the Wikipedia projects. Virtual labour is still conditioned by the diverse working environments across the world, as mentioned by the authors in the quote: below:
"For example, the daily pattern of Asian languages (e.g., Japanese, Chinese and Korean) show higher activity during evenings and nights along with high level of activity at weekends. This can be related partly to the lengths of working hours in corresponding countries. This general image, which holds partially for Turkey and Russia and Israel too, could be in close relation with the high average working hours per day in those countries (more than 40 hours in all the mentioned cases, according to the dataset of *The Organization for Economic Co-operation and Development*: http://stats.oecd.org). Furthermore, among European countries, we also see the same tendency; in the countries with rather larger working times, edits are mostly done in later times in evenings."
Note also that the difference in the timeframes, what I have done by the infographics based on the Wikimedia's Squid reports (not the original traffic data) shows yearly changes. This is in contrast to the Yasseri et al.'s analysis of circadian *daily* and *weekly* patterns. Both have different angles and thus different needs from the Wikimedia Foundation for its traffic data.
Thus, we might want to share what has been done and what could be done regarding the current traffic data provided by the Wikimedia Foundation while acknowledging the sensitivity of the traffic data release,
Best, han-teng liao
2014-05-16 23:50 GMT+08:00 Taha Yasseri taha.yaseri@gmail.com:
Very useful Han-Teng, but one should note that the original data is about "the percentage of requesting ip addresses", excluding duplications of a single IP address within the same day, and not for example the number of edits. These two can be very different depending on dynamic/static IP address models in different countries. And that explains the discrepancy between your results and our earlier analysis based on circadian patterns and edits timestamps. http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0030091#pone-0030091-g004
Again, very interesting and well done. Best, Taha
On Fri, May 16, 2014 at 4:28 PM, h hanteng@gmail.com wrote:
Dear all,
With the aim to compare Wikipedia traffic report data (e.g. viewing versus editing, regional differences within a language version, etc.), I have made a few more interactive infographics which show the historical changes since late 2011. (Historical numbers are scraped from the past versions archived by the Internet archive)
For more, please visit follow the link below: http://people.oii.ox.ac.uk/hanteng/2014/05/16/wikipedia-traffic/
It has at least one nice interactive feature: a user can zoom and pan to view the chart easily with a mouse or mousepad. The SVG vector-based presentation insures the picture quality is consistent when users zoom in to compare data points. (I haven't figured out how mpld3's html tooltip work for this project, though.)
It is also possible to extend the prototype with dynamic jsonobjects so that the chart/tables can be updated automatically.
Any suggestions and comments are welcome.Best, han-teng liao
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
-- .t
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
h, 17/05/2014 01:54:
Thus, we might want to share what has been done
+1, but:
and what could be done regarding the current traffic data provided by the Wikimedia Foundation while acknowledging the sensitivity of the traffic data release,
what additional data would you need and why, given clarifications above? Have you considered using revision data instead, correlated with the publicly available squid reports which already tell you what's the share of each country for each language?
Nemo
Thanks to Federico Leva (Nemo) for the follow up questions (i.e. "what additional data would you need and why, given clarifications above?"). I have the following suggestions for the Wikimedia staff to consider and ask other Wikipedian and Wikipedia researchers to share their thoughts. I organize my answers from the easiest to the most difficult.
Some suggestions to the current and future Wikipedia traffic data curation and presentation.
(I). What can be done quickly with relatively little pain but with substantial gains?
(I-A). Tabulate the data points in absolute numbers first, not percentage numbers In terms of data points, this should be easy to do. It would be much useful if absolute numbers, not percentage numbers, are provided so as to see historical dynamics in terms of absolute numbers not relative numbers in percentage points.
Currently, say for Chinese Wikipedia, it is already possible to see the trend regarding the *proportion* of Taiwan users versus that of Hong Kong. But it would be much more helpful, for researchers and Wikipedians alike, to see if there is a decline or increase in absolute numbers. That is to say, for Chinese Wikipedia, we can know better if, for a specific region, the viewing/editing traffic has increased or decreased.
We need to know whether there is a growth and decline and the current percentage data does not allow us to do so.
(I-B). Include all language versions for the *editing traffic* report as well. In terms of the language version coverage, it would be useful for both Wikipedians and researchers to compare the *editing versus viewing* traffic so as to identify the gaps for development if the editing traffic report would be as comprehensive as that of viewing traffic.
Currently, many language versions are reported with viewing traffic data only, not with editing traffic data. Hindi, Kurdish, Uyghur, Wuu, Cantonese, etc. are such examples: http://users.ox.ac.uk/~kebl3178/wikipedia_traffic_hi.html http://users.ox.ac.uk/~kebl3178/wikipedia_traffic_ku.html http://users.ox.ac.uk/~kebl3178/wikipedia_traffic_ug.html http://users.ox.ac.uk/~kebl3178/wikipedia_traffic_wuu.html http://users.ox.ac.uk/~kebl3178/wikipedia_traffic_zh-yue.html
(I-C). Provide static data objects in more accessible format (i.e. csv and/or json). My life would be much easier if csv and/or json formats are provided. I believe that others would be easier too. For the current outcome, I had to scrape the data off the html page, which was a lot of work.
(II).What should be done soon to provide more consistent and accessible traffic data reports?
(II-A). Putting viewing traffic and editing traffic report on the same page. For viewers' convenience, table presentation and visualization should allow readers to compare editing traffic and viewing traffic *on the same page* so that viewers do not have to switch between pages
(II-B). Organizing and archiving the traffic reports for historical comparison. In terms of traffic data report release (per language and per country one), it would be of great help for Wikipedians and researchers alike to organize past reports according to the coverage of the data points (i.e. annually or quarterly or even monthly) and make the historical pages accessible. (Currently I have to retrieve past data via Internet archive.)
(I-C). Provide dynamic data objects in more accessible format (i.e. csv and/or json). It would be awesome to have some API developed to generate traffic data report in csv or json formats. Note that the infographics that I have prototyped can be tweaked in a way to load data for more interactive experience.
(III).What should be discussed for the longer-term development to inform Wikipedia policies and strategies using/curating the traffic reports? (III-A). Shorter time aggregate units. I notice that there seems to be a shift from providing annual report ones to quarterly ones lately. It is a good direction for others can do the annual average themselves based on quarterly ones.
From researchers' point of view, I would prefer more frequent and shorter
data release cycles (e.g. monthly if not weekly), and then do the statistics (average, etc.) myself so as to derive annual report.
(III-B). Smaller (i.e more specific) geographic aggregate units. The country (geographic) information is often based on geo-IP databases, and sometimes provincial and city-level data would be available. It would be extremely useful if the aggregate units can be lowered one level down to the first administrative levels below countries.
This will create important reports for the geographic distribution of editing/viewing traffic across different provinces in mainland China or India, or different states in the United States.
(III-C). Relevant geolinguistic and geocultural database for country/language name/code disambiguation and queries. First, the country codes and language codes should be provided and maintained centrally in one place so as to help others to reuse the data with data consistency and integrity.
Second, the country names and language names should be also provided and maintained (preferably from Wikidata) so as to help others to localize the traffic data report in all languages! I somehow believe that traffic reports in different languages will help various Wikimedia's outreach programs, including fund-raising. Effectively the country/language names/codes together will provide an important "translation memory" for identifying/converting/translating country and language names/codes.
( I know that the Unicode Common Locale Data Repository (CLDR Version 25http://cldr.unicode.org/index/downloads/cldr-25 ) provides “language-territory” http://www.unicode.org/cldr/charts/latest/supplemental/language_territory_information.html or “territory-language” http://www.unicode.org/cldr/charts/latest/supplemental/territory_language_information.htmlunit-based charts, but I believe that the Wikimedia projects can use and build one better..)
The above suggestions are limited by my own experience and understanding of Wikipedia and content localization/language industry. It is of course biased towards my own interpretations of geolinguistic methods. There are of course other suggestions worthy of considerations and discussions regarding reporting viewing/editing traffic data. I would argue, however, that the geolinguistic comparisons inform a more geocultural (and possibly geopolitical) understanding that also matches how Wikipedia projects are currently divided and governed (language versions with some regional considerations).
Best,
han-teng liao
2014-05-17 14:50 GMT+08:00 Federico Leva (Nemo) nemowiki@gmail.com:
h, 17/05/2014 01:54:
Thus, we might want to share what has been done
+1, but:
and what could be done
regarding the current traffic data provided by the Wikimedia Foundation while acknowledging the sensitivity of the traffic data release,
what additional data would you need and why, given clarifications above? Have you considered using revision data instead, correlated with the publicly available squid reports which already tell you what's the share of each country for each language?
Nemo
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
Thanks for your suggestions. Just some quick pointers below.
h, 18/05/2014 08:26:
(I-A). Tabulate the data points in absolute numbers first, not percentage numbers [...] (I-B). Include all language versions for the *editing traffic* report as well. [...] (I-C). Provide static data objects in more accessible format (i.e. csv and/or json). [...] (II-A). Putting viewing traffic and editing traffic report on the same page. [...] (II-B). Organizing and archiving the traffic reports for historical comparison. [...] (I-C). Provide dynamic data objects in more accessible format (i.e. csv and/or json).
At least the first four are "just" changes in the WikiStats reports formatting, personally I encourage you to submit patches: https://git.wikimedia.org/summary/analytics%2Fwikistats.git (should be the "squids" directory, but there is some ongoing refactoring of the repos).
On archives and "history rewriting"/reports regeneration, see also https://bugzilla.wikimedia.org/show_bug.cgi?id=46198
[...] (III-B). Smaller (i.e more specific) geographic aggregate units. The country (geographic) information is often based on geo-IP databases, and sometimes provincial and city-level data would be available.
http://lists.wikimedia.org/pipermail/wikitech-l/2014-April/075964.html
[...]
( I know that the Unicode Common Locale Data Repository (CLDR Version 25 http://cldr.unicode.org/index/downloads/cldr-25) provides“language-territory” http://www.unicode.org/cldr/charts/latest/supplemental/language_territory_information.htmlor “territory-language” http://www.unicode.org/cldr/charts/latest/supplemental/territory_language_information.htmlunit-based charts, but I believe that the Wikimedia projects can use and build one better..) [...]
No, we definitely can't, not alone. I've asked for help, please contribute: https://www.mediawiki.org/wiki/Universal_Language_Selector/FAQ#How_does_Universal_Language_Selector_determine_which_languages_I_may_understand.
Nemo
Dear Nemo,
As I am waiting for a more complete response, I am not sure that I understand your last "No" as in "No, we definitely can't" means. To clarify, take the CLDR supplement Language-Territory information for example http://www.unicode.org/cldr/charts/latest/supplemental/language_territory_in...
One can suggest additions of the data point by submitting sourced numbers for a geo-linguistic population like this: http://unicode.org/cldr/trac/newticket?&description=%3Cterritory%2c%20sp...)
In Wikipedia articles and Wikidata pages, there are many attempts to provide more updated and better sourced data points. I see the potentials in exchanging such data, curating them better in Wikidata projects as more detailed and dynamic source than the CLDR.
These data points will have extra benefits in curating traffic data. For one, these geo-linguistic population data points would be useful to normalize traffic data for further analysis, such as geographic normalization. For another, they provide important reference data for the development strategies and policies of the Wikipedia projects.
Best, han-teng liao
2014-05-18 16:23 GMT+08:00 Federico Leva (Nemo) nemowiki@gmail.com:
Thanks for your suggestions. Just some quick pointers below.
h, 18/05/2014 08:26:
(I-A). Tabulate the data points in absolute numbers first, not percentage numbers [...]
(I-B). Include all language versions for the *editing traffic* report as well. [...]
(I-C). Provide static data objects in more accessible format (i.e. csv and/or json). [...]
(II-A). Putting viewing traffic and editing traffic report on the same page. [...]
(II-B). Organizing and archiving the traffic reports for historical comparison. [...]
(I-C). Provide dynamic data objects in more accessible format (i.e. csv and/or json).
At least the first four are "just" changes in the WikiStats reports formatting, personally I encourage you to submit patches: < https://git.wikimedia.org/summary/analytics%2Fwikistats.git%3E (should be the "squids" directory, but there is some ongoing refactoring of the repos).
On archives and "history rewriting"/reports regeneration, see also https://bugzilla.wikimedia.org/show_bug.cgi?id=46198
[...] (III-B). Smaller (i.e more specific) geographic aggregate units.
The country (geographic) information is often based on geo-IP databases, and sometimes provincial and city-level data would be available.
http://lists.wikimedia.org/pipermail/wikitech-l/2014-April/075964.html
[...]
( I know that the Unicode Common Locale Data Repository (CLDR Version 25 http://cldr.unicode.org/index/downloads/cldr-25) provides“language-territory” http://www.unicode.org/cldr/charts/latest/supplemental/ language_territory_information.htmlor “territory-language” http://www.unicode.org/cldr/charts/latest/supplemental/ territory_language_information.htmlunit-based
charts, but I believe that the Wikimedia projects can use and build one better..) [...]
No, we definitely can't, not alone. I've asked for help, please contribute: https://www.mediawiki.org/wiki/Universal_Language_ Selector/FAQ#How_does_Universal_Language_Selector_ determine_which_languages_I_may_understand.
Nemo
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
Could you give an example of what we could do better than CLDR or the relevant ISO standards?
On 18 May 2014 10:06, h hanteng@gmail.com wrote:
Dear Nemo,
As I am waiting for a more complete response, I am not sure that Iunderstand your last "No" as in "No, we definitely can't" means. To clarify, take the CLDR supplement Language-Territory information for example
http://www.unicode.org/cldr/charts/latest/supplemental/language_territory_in...
One can suggest additions of the data point by submitting sourcednumbers for a geo-linguistic population like this: http://unicode.org/cldr/trac/newticket?&description=%3Cterritory%2c%20sp...)
In Wikipedia articles and Wikidata pages, there are many attempts toprovide more updated and better sourced data points. I see the potentials in exchanging such data, curating them better in Wikidata projects as more detailed and dynamic source than the CLDR.
These data points will have extra benefits in curating traffic data.For one, these geo-linguistic population data points would be useful to normalize traffic data for further analysis, such as geographic normalization. For another, they provide important reference data for the development strategies and policies of the Wikipedia projects.
Best, han-teng liao
2014-05-18 16:23 GMT+08:00 Federico Leva (Nemo) nemowiki@gmail.com:
Thanks for your suggestions. Just some quick pointers below.
h, 18/05/2014 08:26:
(I-A). Tabulate the data points in absolute numbers first, not percentage numbers [...]
(I-B). Include all language versions for the *editing traffic* report as well. [...]
(I-C). Provide static data objects in more accessible format (i.e. csv and/or json). [...]
(II-A). Putting viewing traffic and editing traffic report on the same page. [...]
(II-B). Organizing and archiving the traffic reports for historical comparison. [...]
(I-C). Provide dynamic data objects in more accessible format (i.e. csv and/or json).
At least the first four are "just" changes in the WikiStats reports formatting, personally I encourage you to submit patches: < https://git.wikimedia.org/summary/analytics%2Fwikistats.git%3E (should be the "squids" directory, but there is some ongoing refactoring of the repos).
On archives and "history rewriting"/reports regeneration, see also https://bugzilla.wikimedia.org/show_bug.cgi?id=46198
[...] (III-B). Smaller (i.e more specific) geographic aggregate units.
The country (geographic) information is often based on geo-IP databases, and sometimes provincial and city-level data would be available.
http://lists.wikimedia.org/pipermail/wikitech-l/2014-April/075964.html
[...]
( I know that the Unicode Common Locale Data Repository (CLDR Version 25 http://cldr.unicode.org/index/downloads/cldr-25) provides“language-territory” http://www.unicode.org/cldr/charts/latest/supplemental/ language_territory_information.htmlor “territory-language” http://www.unicode.org/cldr/charts/latest/supplemental/ territory_language_information.htmlunit-based
charts, but I believe that the Wikimedia projects can use and build one better..) [...]
No, we definitely can't, not alone. I've asked for help, please contribute: https://www.mediawiki.org/wiki/Universal_Language_ Selector/FAQ#How_does_Universal_Language_Selector_ determine_which_languages_I_may_understand.
Nemo
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
Hello Oliver, Let me use Cantonese (yue) and Hakka (hak) as examples to illustrate some possibilities. Just the population data points. Have a look at the Ethnologue data http://www.ethnologue.com/language/yueand http://www.ethnologue.com/language/hak Note that you should see the population number in China and also other places in the world (under the section of "Also Spoken In") There are also other data points such as "status" and "writing". Then one can look up the CLDR's Language-Territory or Territory-Language information, the entries for Cantonese and Hakka does not exist yet. Note also that both Cantonese and Hakka have their own language versions of Wikipedia (zh-yue and hak). The coding and naming needs a table here for data integration. Now, as tertiary sources that integrates other data points, Wikipedia/Wikidata can get the data points from Ethnologue to enrich its content. These data points would be important baseline for almost any human language-based Wikipedia projects to identify their potential editors. The current active editors of small and medium size language Wikipedia projects should be interested in getting hold of such data. Also, they may know more updated and reliable data ahead of Ethnologue. For traffic data reports, a Cantonese Wikipedian can then normalize the viewing and editing traffic data against the population data, thereby identifying the "per speaker capita" number for the viewing/editing traffic. I have done some normalization work (or geolinguistic normalization) for languages such as Spanish and Arabic where the CLDR's Language-Territory or Territory-Language information data. The surprising results are that for Spanish, per captia editing traffic are the highest in Germany, Paraguay, Uruguay and Spain; per capita viewing traffic are the highest in Paraguay, Spain, Chile, etc. For Arabic, per capita editing traffic are the highest in Kuwait, Baharain, Saudi Arabia, Qatar, Israel, UAE, etc; per capita viewing traffic are the highest in Israel, Kuwait, Saudi Arabia, etc. I personally believe such data curation, when supported by better and expected-to-be-improved geolinguistic data population data points now available in Ethnologue and other sources that different language Wikipedians may know, would be useful to Wikipedians first. In short, I did not intend to ask Wikipedians or the Wikimedia research staff to do extra "original research". My suggestions aim to parse the traffic data one level down from either language or territory to the more specific language-territory aggregate so as better inform development strategies and academic research on Wikipedia. Overall, I think it is viable to construct a data process to show what need to be done and what can be achieved. The showing-by-doing approach can show some results first with infographics for language versions that are more data-ready (e.g. Arabic and Spanish). Then other language versions can strive to fill the now *identified* data gaps by contributing data points through Wikipedia and Wikidata projects. What is needed then is a database and expert pool of territory-language and language-territory information across Wikipedia projects. It can be as simple and as straightforward to have a Wikidata object of geo-lingustic population for any territory-language combinations, potentially with existing translations made possible by Wikidata, then the traffic/viewing data reports can be (1) localized/translated into different languages automatically and (2) geo-linguistically normalized to show the current outreach of a language Wikipedia per language-speaker. The above are only my current rough and initial thoughts. Please let me know if the ideas or expressions are not clear enough. Best, han-teng liao
2014-05-19 7:35 GMT+08:00 Oliver Keyes okeyes@wikimedia.org:
Could you give an example of what we could do better than CLDR or the relevant ISO standards?
On 18 May 2014 10:06, h hanteng@gmail.com wrote:
Dear Nemo,
As I am waiting for a more complete response, I am not sure that Iunderstand your last "No" as in "No, we definitely can't" means. To clarify, take the CLDR supplement Language-Territory information for example
http://www.unicode.org/cldr/charts/latest/supplemental/language_territory_in...
One can suggest additions of the data point by submitting sourcednumbers for a geo-linguistic population like this: http://unicode.org/cldr/trac/newticket?&description=%3Cterritory%2c%20sp...)
In Wikipedia articles and Wikidata pages, there are many attempts toprovide more updated and better sourced data points. I see the potentials in exchanging such data, curating them better in Wikidata projects as more detailed and dynamic source than the CLDR.
These data points will have extra benefits in curating traffic data.For one, these geo-linguistic population data points would be useful to normalize traffic data for further analysis, such as geographic normalization. For another, they provide important reference data for the development strategies and policies of the Wikipedia projects.
Best, han-teng liao
2014-05-18 16:23 GMT+08:00 Federico Leva (Nemo) nemowiki@gmail.com:
Thanks for your suggestions. Just some quick pointers below.
h, 18/05/2014 08:26:
(I-A). Tabulate the data points in absolute numbers first, not percentage numbers [...]
(I-B). Include all language versions for the *editing traffic* report as well. [...]
(I-C). Provide static data objects in more accessible format (i.e. csv and/or json). [...]
(II-A). Putting viewing traffic and editing traffic report on the same page. [...]
(II-B). Organizing and archiving the traffic reports for historical comparison. [...]
(I-C). Provide dynamic data objects in more accessible format (i.e. csv and/or json).
At least the first four are "just" changes in the WikiStats reports formatting, personally I encourage you to submit patches: < https://git.wikimedia.org/summary/analytics%2Fwikistats.git%3E (should be the "squids" directory, but there is some ongoing refactoring of the repos).
On archives and "history rewriting"/reports regeneration, see also https://bugzilla.wikimedia.org/show_bug.cgi?id=46198
[...] (III-B). Smaller (i.e more specific) geographic aggregate units.
The country (geographic) information is often based on geo-IP databases, and sometimes provincial and city-level data would be available.
http://lists.wikimedia.org/pipermail/wikitech-l/2014-April/075964.html
[...]
( I know that the Unicode Common Locale Data Repository (CLDR Version 25 http://cldr.unicode.org/index/downloads/cldr-25) provides“language-territory” http://www.unicode.org/cldr/charts/latest/supplemental/ language_territory_information.htmlor “territory-language” http://www.unicode.org/cldr/charts/latest/supplemental/ territory_language_information.htmlunit-based
charts, but I believe that the Wikimedia projects can use and build one better..) [...]
No, we definitely can't, not alone. I've asked for help, please contribute: https://www.mediawiki.org/wiki/Universal_Language_ Selector/FAQ#How_does_Universal_Language_Selector_ determine_which_languages_I_may_understand.
Nemo
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
-- Oliver Keyes Research Analyst Wikimedia Foundation
Wiki-research-l mailing list Wiki-research-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wiki-research-l
wiki-research-l@lists.wikimedia.org