Just as a suggestion, you can turn these kind of numbers into a probability distribution using the beta distribution. If you use (1,1) as a prior you get something like beta(251,1) for the the probability of the probability that somebody named "Aaron" is male.
-----Original Message----- From: Markus Krötzsch Sent: Sunday, October 13, 2013 6:16 PM To: Discussion list for the Wikidata project. Subject: [Wikidata-l] Application: sexing people by name/research gender bias
Hi all,
I'd like to share a little Wikidata application: I just used Wikidata to guess the sex of people based on their (first) name [1]. My goal was to determine gender bias among the authors in several research areas. This is how some people spend their free time on weekends ;-)
In the process, I also created a long list of first names with associated sex information from Wikidata [2]. It is not super clean but it served its purpose. If you are a researcher, then maybe the gender bias of journals/conferences is interesting to you as well. Details and some discussion of the results are online [1].
Cheers,
Markus
[1] http://korrekt.org/page/Note:Sex_Distributions_in_Research [2] https://docs.google.com/spreadsheet/ccc?key=0AstQ5xfO-xXGdE9UVkxNc0JMVWJzNmJ...
_______________________________________________ Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
Hi all,
First of all I think this is fantastic research. It goes to show, it's not just properties that we can correlate, but also the Labels, Aliases, Sitelinks, and the connections between each field.
I would like to point out, as Markus does in his discussion - the relative disproportionate representation of sex in Acadmia is the motivation for studying this. Let us be sensitive to results in that field. Lets remember our simplifying assumptions. We have flattened sex and gender into one measure, and at that this research makes a binary male/female classification, where even the wikidata sex property is trinary (intersex). I hope that in the future we can increase or change our view to how we model sex.
Best,
Maximilian Klein Wikipedian in Residence, OCLC +17074787023
________________________________________ From: wikidata-l-bounces@lists.wikimedia.org wikidata-l-bounces@lists.wikimedia.org on behalf of Paul A. Houle paul@ontology2.com Sent: Sunday, October 13, 2013 5:32 PM To: Discussion list for the Wikidata project. Subject: Re: [Wikidata-l] Application: sexing people by name/research gender bias
Just as a suggestion, you can turn these kind of numbers into a probability distribution using the beta distribution. If you use (1,1) as a prior you get something like beta(251,1) for the the probability of the probability that somebody named "Aaron" is male.
-----Original Message----- From: Markus Krötzsch Sent: Sunday, October 13, 2013 6:16 PM To: Discussion list for the Wikidata project. Subject: [Wikidata-l] Application: sexing people by name/research gender bias
Hi all,
I'd like to share a little Wikidata application: I just used Wikidata to guess the sex of people based on their (first) name [1]. My goal was to determine gender bias among the authors in several research areas. This is how some people spend their free time on weekends ;-)
In the process, I also created a long list of first names with associated sex information from Wikidata [2]. It is not super clean but it served its purpose. If you are a researcher, then maybe the gender bias of journals/conferences is interesting to you as well. Details and some discussion of the results are online [1].
Cheers,
Markus
[1] http://korrekt.org/page/Note:Sex_Distributions_in_Research [2] https://docs.google.com/spreadsheet/ccc?key=0AstQ5xfO-xXGdE9UVkxNc0JMVWJzNmJ...
_______________________________________________ Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
_______________________________________________ Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
On 14/10/13 17:52, Klein,Max wrote:
Hi all,
First of all I think this is fantastic research. It goes to show, it's not just properties that we can correlate, but also the Labels, Aliases, Sitelinks, and the connections between each field.
I would like to point out, as Markus does in his discussion - the relative disproportionate representation of sex in Acadmia is the motivation for studying this. Let us be sensitive to results in that field. Lets remember our simplifying assumptions. We have flattened sex and gender into one measure, and at that this research makes a binary male/female classification, where even the wikidata sex property is trinary (intersex). I hope that in the future we can increase or change our view to how we model sex.
Indeed, it the debates on "gender inequality" and "gender multiplicity" look at things on very different zoom levels. The goal of my little experiment (I would not call it research, as it has neither a hypothesis nor any form of evaluation) was not to put individual people into rigid gender buckets but to estimate rough global distributions. My error margins are far too wide to make any realistic statement about "minority genders" even if I had a method to consider them. As far as social definitions of gender go, this is probably something to study in a wider context of representation of social minorities in certain professional fields.
Cheers,
Markus
From: wikidata-l-bounces@lists.wikimedia.org wikidata-l-bounces@lists.wikimedia.org on behalf of Paul A. Houle paul@ontology2.com Sent: Sunday, October 13, 2013 5:32 PM To: Discussion list for the Wikidata project. Subject: Re: [Wikidata-l] Application: sexing people by name/research gender bias
Just as a suggestion, you can turn these kind of numbers into a probability distribution using the beta distribution. If you use (1,1) as a prior you get something like beta(251,1) for the the probability of the probability that somebody named "Aaron" is male.
-----Original Message----- From: Markus Krötzsch Sent: Sunday, October 13, 2013 6:16 PM To: Discussion list for the Wikidata project. Subject: [Wikidata-l] Application: sexing people by name/research gender bias
Hi all,
I'd like to share a little Wikidata application: I just used Wikidata to guess the sex of people based on their (first) name [1]. My goal was to determine gender bias among the authors in several research areas. This is how some people spend their free time on weekends ;-)
In the process, I also created a long list of first names with associated sex information from Wikidata [2]. It is not super clean but it served its purpose. If you are a researcher, then maybe the gender bias of journals/conferences is interesting to you as well. Details and some discussion of the results are online [1].
Cheers,
Markus
[1] http://korrekt.org/page/Note:Sex_Distributions_in_Research [2] https://docs.google.com/spreadsheet/ccc?key=0AstQ5xfO-xXGdE9UVkxNc0JMVWJzNmJ...
Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
On Tue, Oct 15, 2013 at 7:50 AM, Markus Krötzsch < markus@semantic-mediawiki.org> wrote:
My error margins are far too wide to make any realistic statement about "minority genders" even if I had a method to consider them.
This article: http://journal.code4lib.org/articles/8964 gives them as being in the range 0.002% - 0.006% so they're unlikely to effect any real-world analysis.
Tom
Hi Markus and Tom,
Markus, I understand that your report does not claim to be research. I actually happen to think it's a really fantastic and mind-expanding idea of what we can do with Wikidata. It's given me much food for thought, not only can we correlate Wikidata fields, but we can use that data as predictors. That's a step further than I ever imagined to go. What I wrote, and am writing, is not intended to be a criticism of your efforts.
Tom, the link you are citing is my own paper, and the data I reported is purely "positive" and "descriptive" not only Wikipedia, but also the subset of Wikipedia data that has been migrated to Wikidata. Which is to say that it contains all the biases of those processes. That is not necessarily bad - but what might be is the reasoning that stems from interpreting these results. For instance, depending on how you define sexually ambiguous humans, between 0.1% and 1.7% of humans could be classified as such. [1] So conservatively Wikidata is two orders of magnitude off, and could be three orders of magnitude off. And then all of a sudden were interpreting our own bias as truth, and possibly just simplifying people out of existence.
I'm not saying anyone has done anything wrong. I just feel abstractly concerned - that's nobody's fault in particular. Somehow Wikidata has given us the power to greater quantify our view of the world, and our bias is really becoming clear - numerically. Then we think about things like make a bot to give people properties based on strings, and placing value constraints on the sex property. Who is that helping? The software is beautifully built so it doesn't force us to do any of this. I would argue it's our inherited worldviews that is guiding us.
Sorry to rant. These are my feelings, they do not require a response. I do not demand or request that anybody feel the same way.
[1] https://en.wikipedia.org/wiki/Intersex#Prevalence. (Citing the underlying citations of https://www.worldcat.org/search?qt=wikipedia&q=isbn%3A0465077137)
Maximilian Klein Wikipedian in Residence, OCLC +17074787023
________________________________ From: wikidata-l-bounces@lists.wikimedia.org wikidata-l-bounces@lists.wikimedia.org on behalf of Tom Morris tfmorris@gmail.com Sent: Tuesday, October 15, 2013 9:14 AM To: Discussion list for the Wikidata project. Subject: Re: [Wikidata-l] Application: sexing people by name/research gender bias
On Tue, Oct 15, 2013 at 7:50 AM, Markus Kr?tzsch <markus@semantic-mediawiki.orgmailto:markus@semantic-mediawiki.org> wrote: My error margins are far too wide to make any realistic statement about "minority genders" even if I had a method to consider them.
This article: http://journal.code4lib.org/articles/8964 gives them as being in the range 0.002% - 0.006% so they're unlikely to effect any real-world analysis.
Tom
So you've got an agenda that's unrelated to Wikidata or analysis thereof. Got it. Perhaps a non-Wikidata list would be a more appropriate forum.
On Tue, Oct 15, 2013 at 2:08 PM, Klein,Max kleinm@oclc.org wrote:
Sorry to rant.
Accepted.
Tom
I think the results of Max are really interesting and fruitful, and should be shared with this list here: https://lists.wikimedia.org/mailman/listinfo/gendergap
Aubrey
On Tue, Oct 15, 2013 at 8:33 PM, Tom Morris tfmorris@gmail.com wrote:
So you've got an agenda that's unrelated to Wikidata or analysis thereof. Got it. Perhaps a non-Wikidata list would be a more appropriate forum.
On Tue, Oct 15, 2013 at 2:08 PM, Klein,Max kleinm@oclc.org wrote:
Sorry to rant.
Accepted.
Tom
Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
Max's comment is very related to Wikidata. The sex property [1] is a model system to explore important questions for the project at large.
For example, how rigorous do we want to be with automatic classification? Let's say a property can have one of three values: A, B or C. Roughly 90% of the valid subjects for that property are known to be either A or B, and 10% are known to be C. Our automatic classifier can assign all valid subjects to either A or B. However, it can't segregate A or B from C. So our false positive rate is at least 10%. Would it be acceptable for Wikidata to have a known error rate of 10% in certain properties? At what error rate does automatic classification become unacceptable?
Another question this topic broaches: do we want to adopt formal domain and range constraints on properties? If we do, then how do we handle rare values? How about exceedingly rare values? (It should be noted that the Wikidata sex property includes intersex in its range constraints [2].) There is ongoing discussion about whether we want to adopt range and domain constraints (among other property metadata) in Wikidata's Project chat [3].
Eric https://www.wikidata.org/wiki/User:Emw
1. https://www.wikidata.org/wiki/Property:P21 2. https://www.wikidata.org/wiki/Property_talk:P21 3. https://www.wikidata.org/wiki/Wikidata:Project_chat#What_type_of_data_should...: https://www.wikidata.org/w/index.php?title=Wikidata:Project_chat&oldid=7... )
On Tue, Oct 15, 2013 at 2:33 PM, Tom Morris tfmorris@gmail.com wrote:
So you've got an agenda that's unrelated to Wikidata or analysis thereof. Got it. Perhaps a non-Wikidata list would be a more appropriate forum.
On Tue, Oct 15, 2013 at 2:08 PM, Klein,Max kleinm@oclc.org wrote:
Sorry to rant.
Accepted.
Tom
Wikidata-l mailing list Wikidata-l@lists.wikimedia.org https://lists.wikimedia.org/mailman/listinfo/wikidata-l
On 15/10/13 19:08, Klein,Max wrote: ...
I'm not saying anyone has done anything wrong. I just feel abstractly concerned - that's nobody's fault in particular. Somehow Wikidata has given us the power to greater quantify our view of the world, and our bias is really becoming clear - numerically.
+1 to that. For example, I noticed from Tom's post that Freebase knows about 500 people called Nicola [1], while I found only 320 on Wikidata:
* Freebase: 248 men and 257 men * Wikidata: 276 men and 50 women
Who are these mysterious 200+ female Nicolas are that Freebase cares about while Wikipedia doesn't? The general size of the numbers suggests that both datasets are highly selective, considering only a tiny percentage of the world's Nicolas to be notable enough to be included. (Of course, we might also find that their men are not the same as ours.)
Markus
[1] http://namegender.freebaseapps.com/gender_api?name=nicola