[Flow Funding] Potential project

Andrew G. West westand at cis.upenn.edu
Tue Mar 5 19:33:28 UTC 2013


Hi flow funders,

I have spent a good bit of time recently building and advertising my 
flow funding portal. Some of you might be able to use this as a model:

http://en.wikipedia.org/wiki/User:West.andrew.g/Flow_funding

You may also notice that it recently attracted its first proposal -- and 
I have received a second one via email [[attached]] that needs to be 
transferred into wiki-space.

I would welcome feedback on either of these (and I am currently quite 
inclined to support the one on the wiki).

I think a tricky point in the funding language regards the notion that 
people should not be compensated for labor "that would normally be 
performed by volunteer actions." Where does one draw this line?

Thanks, -AW


On 02/26/2013 08:54 PM, Kiril Simeonovski wrote:
> Dear flow funders,
>
> A professor at the Faculty of Economics at Skopje University is
> interested to implement modern techniques in his curriculum by
> introducing the Wikimedia projects, lectures related to some aspects of
> the Wikis, and invite speakers from other universities in the region. I
> still don't have much information about the budget for this project, but
> I was told that it won't be costly and therefore decided to mediate in
> the process about the project. The basic idea is to invite a special
> guest from abroad who will have to give several lectures in specific
> subjects of statistics to the auditorium at the faculty. Meanwhile, the
> students will be assigned to create content to the Wikimedia projects
> (specifically Wikipedia, Wikibooks and possibly Wikiversity) and post it
> until a previously defined deadline. One of the main roles of the guest
> is to be his commitment to link to other universities abroad and propose
> the implementation of a project in a similar way. The latter can be
> easily processed through the existing Wikimedia chapter or community in
> the home country of the guest-speaker.
>
> When I heard for the idea about this project, it sounded pretty good and
> interesting to me. Can you please share some of your thoughts whether
> it's viable to support a such project through the flow funding process?
>
> Best regards,
>
> Kiril Simeonovski
>
>
> _______________________________________________
> FlowFunding mailing list
> FlowFunding at lists.wikimedia.org
> https://lists.wikimedia.org/mailman/listinfo/flowfunding
>

-- 
Andrew G. West, Doctoral Candidate
Dept. of Computer and Information Science
University of Pennsylvania, Philadelphia PA
Email:   westand at cis.upenn.edu
Website: http://www.andrew-g-west.com
-------------- next part --------------

Username(s) and affiliation(s): Who are you? What projects do you participate in? Any real-life institutions that are relevant?
=============

My name is Andrew Krizhanovsky. 
I am researcher in Institute of Applied Mathematical Research of the Karelian Research Centre of the Russian Academy of Sciences (IAMR).
My project devoted to extraction of information from Wiktionary.
Project page: 
http://code.google.com/p/wikokit/

I) The maximum goal (in distant future) is to extract all information (i.e. all sections of entry) from all wiktionaries and convert data to machine-readable database.

II) Today's result. Now machine-readable Wiktionary contains the following information extracted from Russian Wiktionary and English Wiktionary:

    word's language and part of speech;
    meanings / definitions;
    semantic relations;
    translations. 

Also the quotations were extracted from the Russian Wiktionary. 

III) The next step (supported by grant, I hope) will be an extraction and analysis of "context labels" from Russian and English Wiktionaries.
See http://en.wiktionary.org/wiki/Wiktionary:Context_labels

"Context labels" are very important lexicographic information, which could be used in many tasks of Computational linguistics (e.g. "sentiment analysis", "word sense disambiguation").

There is an analog of context labels in Wordnet. It is Wordnet Domains, 
see http://scholar.google.com/scholar?q=%22wordnet+domain%22&btnG=&hl=ru&as_sdt=0

It will be interesting in the research paper to compare WordNet Domains with Wiktionary context labels.


Participant project roles: What does your project/research history look like? Why should I believe you can complete this work?
=============
I am researcher and programmer. See about me: http://en.wikipedia.org/wiki/User:AKA_MBG

Results of my research (open-source software and publications) are available online. 
See list of my publications related to Wiktionary: http://code.google.com/p/wikokit/#Further_reading

Last two years my research was devoted to the quantitative analysis of Russian and English Wiktionaries and comparison with WordNet thesaurus. During this quantitative linguistics research the number of words in Wiktionary, the distribution of words for each part of speech, the quantity of monosemous and polysemous words, etc. were analysed. See publication: http://scipeople.com/publication/108159/


Concept abstract: What do you want to do? Why? What is the eventual output? How long will it take?
=============
Goal: To extract and analyse "context labels" of Russian and English Wiktionaries.

Results: 
1) Source code of the Wiktionary parser (http://code.google.com/p/wikokit) will be extended by the module to extract "context labels" from Russian Wiktionary and English Wiktionary. 
2) The machine-readable Wiktionary database will be extended by "context labels" extracted from Russian Wiktionary and English Wiktionary. These two constructed databases will be available online at http://code.google.com/p/wikokit/downloads/list
3) The research results will be available in the form of online preprint. In this paper the quantitative analysis and comparison of "context labels" in Russian Wiktionary and English Wiktionary will be carry out.

It will be interesting research, because there are several hundreds of different labels in both Wiktionaries, and they are not always correspond each other. 
See labels in English Wiktionary:
http://en.wiktionary.org/wiki/Category:Context_labels

See labels in Russian Wiktionary:
http://ru.wiktionary.org/wiki/%D0%9A%D0%B0%D1%82%D0%B5%D0%B3%D0%BE%D1%80%D0%B8%D1%8F:%D0%A8%D0%B0%D0%B1%D0%BB%D0%BE%D0%BD%D1%8B_%D0%BF%D0%BE%D0%BC%D0%B5%D1%82

So, the first task of this research is to make an alignment of these labels in two Wiktionaries.

Impact on WMF initiatives: How does this help the project(s)? What is the expected impact/scope?
=============
1) The comparison and analysis of Wiktionaries will help editors to understand:
- gaps and shortcomings in context labels system (in comparison with other Wiktionary);
- which labels are frequently used and which labels are rare used?;
- list of frequent combinations of context labels (label pattern) used in order to describe one meaning in the Wiktionary entry.
- how many entries in the Wiktionary with labels?

2) The open-source machine-readable Wiktionary extended by context labels could be used in oder to solve many computational linguistics tasks.

3) The developed free Android applications (kiwidict and kiwidict-ru) are based on the machine-readable Wiktionary. So databases of these applications will be extended by context labels. These applications provide unique search possibilities which is not available at the Wiktionary site. E.g. you can search words in one (or several) language using wildcard search (% and _ symbols).


Amount requested & budgetary justification: How much do you want? Why is that amount appropriate? Why are cheaper/free alternatives inappropriate?
=============
2000$ is fine, 1000$ is OK also :)
This project will require 6-9 month. And it will be finished in this year (2013).
This money will be spend on my salary during this project.



More information about the FlowFunding mailing list