Forwarding some job listings, including selected quotes from the listings.
1. *https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/post_doctoral_researcher_15072
<https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/post_doctora…>*
- applications due 8 April. "The Max Planck Institute for Demographic
Research (MPIDR) is seeking to appoint a full-time post-doctoral researcher
to join the Research Group on Labor Demography." "We welcome applications
from researchers with a PhD in demography, statistics, sociology,
economics, epidemiology, public health, or allied fields." "Strong
quantitative analysis skills are required, as is knowledge of R, Stata, or
Python."
2. *https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/max_planck_postdoc_positions_on_artefacts_for_the_social_sciences_15266
<https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/max_planck_p…>*
-
applications due 13 April. "The Max Planck Institute for Software Systems
(Kaiserslautern and Saarbruecken), the Max Planck Institute for Security
and Privacy (Bochum), The Max Planck Institute for Demographic Research
(Rostock), and the Max Planck Institute for the Study of Crime, Security
and Law (Freiburg) are jointly recruiting up to 8 highly qualified
Post-Docs interested in Computational Social Science. This application call
is targeted to the broad topic of artificial intelligence, computing, and
society."
3. *https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/doctoral_student_position_15039
<https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/doctoral_stu…>*
-
applications due 15 April. "The Max Planck Institute for Demographic
Research (MPIDR) is inviting applications from qualified and highly
motivated students for a doctoral student position in the Lab of Migration
and Mobility of the Department of Digital and Computational Demography."
"The Lab of Migration and Mobility of the Department of Digital and
Computational Demography, headed by MPIDR Director Emilio Zagheni, is
looking for candidates with strong quantitative and programming skills, and
with research interests in the fields of demography, migration studies,
science of science, and computational social science."
Pine🌲
Hello,
The Data Platform SRE team needs to carry out an upgrade of our
*dse-k8s-eqiad* Kubernetes cluster, which will require some downtime for
services hosted on this cluster.
If you don't use any of these services, you can ignore the rest of this
message.
* Airflow
* DataHub
* GrowthBook
* Flink
* Spark History
* Superset
* Test Kitchen
These services will, unfortunately, require downtime while the
Kubernetes upgrade is in progress.
The maintenance window, during which we will carry out the upgrade is
this *Thursday March 26th, between 10:30 and 15:00 UTC*.
Some of our dse-k8s workload is already available in a multi-dc capacity
on the *dse-k8s-codfw* cluster, so will *not* require downtime.
Specifically, these are our OpenSearch on Kubernetes clusters for:
* opensearch-ipoid
* opensearch-semantic-search
We will be keeping this ticket updated throughout the maintenance
window, on our progress:
T414484 Upgrade DSE clusters to kubernetes 1.31
<https://phabricator.wikimedia.org/T414484>
Once the Airflow services are running again, we will be working with
representatives from the Data Engineering team to ensure that all data
pipeline tasks have been addressed.
If you have any queries, comments, or concerns, please don't hesitate to
share them with us.
You can get in touch with us by reply to this message, or via
#data-platform-sre on Slack, or #wikimedia-data-platform on IRC.
Kind regards,
Ben
--
*Ben Tullis*(he/him)
Staff Site Reliability Engineer
Wikimedia Foundation <https://wikimediafoundation.org/>
Has the format of the data lake history files changed? Some code I have which works with these assumes 70 fields, and that jives with what's documented at https://wikitech.wikimedia.org/wiki/Data_Platform/Data_Lake/Edits/MediaWiki…. But it looks like there's now 76 fields:
$ bzcat 2026-01.enwiki.2026-01.tsv.bz2 | head -1 | tr '\t' '\n' | wc -l
76
Hello,
We have to carry out some maintenance of our Druid
<https://wikitech.wikimedia.org/wiki/Data_Platform/Systems/Druid> public
cluster, which will unfortunately require a period of service downtime.
The work is scheduled for next Monday, March 16th at approximately 10:30
UTC and is expected to take between 30 and 60 minutes.
Systems adversely affected will be:
* Superset - when querying the Druid Public (AQS) database
* Analytics Query Services (AQS) APIs:
o
Edit Analytics
o
Editor Analytics
* Global Editor Metrics
o Mobile Apps activity tab
o Growth Team impact module
o Year in review
We will be upgrading the Druid cluster during this time, as well the
operating systems of the hosts that serve the cluster.
Sadly, we are unable to carry out a rolling upgrade of the Druid
cluster, so this downtime is unavoidable.
If you have any particular concerns or queries about this maintenance
window, then please get in touch and we can look into deferring or
rescheduling this work.
Kind regards,
Ben
--
*Ben Tullis*(he/him)
Staff Site Reliability Engineer
Wikimedia Foundation <https://wikimediafoundation.org/>
Hi, I'm forwarding the below to additional email lists in case the
announcement interests additional people who aren't subscribed to
Wiktech-l. I suggest that any questions or comments for this topic which
are intended for public email discussion be placed in the original thread
on Wikitech-l
Regards,
Pine🌲
---------- Forwarded message ---------
From: Jonathan Tweed via Wikitech-l <wikitech-l(a)lists.wikimedia.org>
Date: Mon, Mar 2, 2026 at 8:44 AM
Subject: [Wikitech-l] New global API rate limits
To: <wikitech-l(a)lists.wikimedia.org>
Cc: Jonathan Tweed <jtweed(a)wikimedia.org>
Hi all
To help ensure fair and sustainable access
<https://www.mediawiki.org/wiki/MediaWiki_Product_Insights/Content_Reuse>
to Wikimedia resources, over the next month the Wikimedia Foundation will
implement global API rate limits across our APIs.
Why we’re doing this
As we’ve shared over the last year
<https://diff.wikimedia.org/2025/04/01/how-crawlers-impact-the-operations-of…>,
since 2024 we’ve observed a significant rise in automated requests, across
scraping, APIs and bulk downloads. This is continuing to cause
unsustainable load on our infrastructure, taking time and resources away
that we need to support the Wikimedia projects, contributors and readers.
To ensure we can continue to provide preferential access for human and
mission-oriented traffic, we need to reduce the amount of unidentified
requests to our APIs from its current level of around 33%. This reduction
will also enable us to improve governance around fair use, in line with our API
Usage Guidelines
<https://foundation.wikimedia.org/wiki/Policy:Wikimedia_Foundation_API_Usage…>
.
What we’re doing
In early March, we will apply low limits to anonymous API requests that
originate from outside Toolforge/WMCS and requests that are made from web
browsers. In early April, higher limits will be applied to identified
traffic, including authenticated requests.
Authenticated requests from a user in the ‘bot’ user group on any wiki will
not be subject to these new limits, nor will clients which are well known
to the Wikimedia Foundation. API requests from Toolforge/WMCS will also be
exempt from rate limits for now.
Regardless, all developers are encouraged to familiarise themselves with
the new limits and associated best practices to ensure they are not
misclassified as abusive traffic. You can find this information at Wikimedia
APIs/Rate limits <https://www.mediawiki.org/wiki/Wikimedia_APIs/Rate_limits>
.
What this means in practice
Tools that use Wikimedia-hosted APIs, including Gadgets that make API
requests, may be subject to new cross-API limits that are global across all
Wikimedia projects. These limits are intentionally set as high as possible
to minimise impact on the community.
We are asking developers to identify their requests to access higher
limits. Tools running in Toolforge/WMCS are exempted for now, but we
request that you follow the User-Agent policy
<https://foundation.wikimedia.org/wiki/Policy:Wikimedia_Foundation_User-Agen…>
and provide a meaningful User-Agent to help us correctly identify the
source of traffic.
Otherwise, authenticating using session cookies or OAuth 2.0 will grant a
higher limit. Where the authenticated user has the community approved bot
right on any wiki, this will also exempt you from limits, even when the
tool is not running on Toolforge/WMCS.
This request for authentication is a change from the previous guidance in
the Robot policy, which suggested not authenticating to improve cache hit
rates. We do not expect any significant impact from this change with
current usage patterns and will be working over the next year to improve
caching for authenticated requests.
Community impact
A key principle throughout this work has been to ensure responsible use of
our infrastructure, whilst minimizing the impact on the community. Our
communities rely on a broad ecosystem of bots and other tools to create and
maintain the wikis, created by a dedicated group of technical volunteers.
These limits are being put in place to protect our projects from high
levels of abuse and ensure that we are able to better insulate our
community infrastructure from high-volume commercial usage. They are
necessary to ensure that developers are using the most appropriate
channels, giving us the ability to prioritize the community and our human
readers.
Ideally, members of Wikimedia communities will not be affected by this
change. However, it is possible that a small number of bots and other tools
which operate at a very high rate may get rate limited. We ask you to
follow the best practices and are here to help bot operators get the access
they need.
What we need you to do
If you are a developer that uses Wikimedia APIs, we ask you to:
-
Read more about the new limits
<https://www.mediawiki.org/wiki/Wikimedia_APIs/Rate_limits> and the
updated Robot policy <https://wikitech.wikimedia.org/wiki/Robot_policy>
-
Update your tools/bots to follow the new best practices
Should you require a higher rate of access, you are able to:
-
Request the bot flag <https://www.mediawiki.org/wiki/Manual:Bots> from
your local wiki community
-
Consider running in Toolforge or another WMCS offering
-
Use Wikimedia Enterprise APIs
<https://meta.wikimedia.org/wiki/Wikimedia_Enterprise#Access> for
high-volume usage
-
Contact the WMF at bot-traffic(a)wikimedia.org
Best
Jonathan
--
*Jonathan Tweed* (he/him)
Senior Product Manager, Core Platform
Wikimedia Foundation <https://wikimediafoundation.org/>
_______________________________________________
Wikitech-l mailing list -- wikitech-l(a)lists.wikimedia.org
To unsubscribe send an email to wikitech-l-leave(a)lists.wikimedia.org
https://lists.wikimedia.org/postorius/lists/wikitech-l.lists.wikimedia.org/
Hi everyone,
The February 2026 Research Showcase will be live-streamed this Wednesday,
February 25, at 9:30 AM PT / 17:30 UTC. Find your local time here
<https://zonestamp.toolforge.org/1772040600>. Our theme this month is *AI
and Communities*.
*We invite you to watch via the YouTube
stream: https://www.youtube.com/live/qW5IQJv84HY
<https://www.youtube.com/live/qW5IQJv84HY>.* As always, you can join the
conversation in the YouTube chat as soon as the showcase goes live.
This month, we will have two presentations:
*LLMs in Wikipedia: Investigating How LLMs Impact Participation in
Knowledge Communities*By *Moyan Zhou (University of Minnesota)*Large
language models (LLMs) are reshaping knowledge production as community
members increasingly incorporate them into their contribution workflows.
However, participating in knowledge communities involves more than just
contributing content - it is also a deeply social process shaped by
members' level of expertise. While communities must carefully consider
appropriate and responsible LLM integration, the absence of concrete norms
has left individual editors to experiment and navigate LLM use on their
own. Understanding how LLMs influence community participation across
expertise levels is therefore critical in shaping future norms and
supporting effective adoption. To address this gap, we investigated
Wikipedia, one of the largest knowledge production communities, to
understand participation in three dimensions: 1) how LLMs influence the
ways editors gather knowledge, 2) how editors leverage strategies to align
LLM outputs with community norms, and 3) how other editors in the community
respond to LLM-assisted contributions. Through interviews with 16 Wikipedia
editors with different levels of expertise who had used LLMs for their
edits, we revealed a participation gap mediated by expertise in adopting
LLMs in knowledge contributions across knowledge gathering, alignment with
community norms, and peer responses. Based on these findings, we challenge
existing models of novice editors' involvement and propose design
implications for LLMs that support community engagement, highlighting
opportunities for LLMs to sustain mentorship, knowledge transmission, and
legitimacy building by scaffolding and feedback, process documentation, and
LLM disclosure by good-faith editors.*AI Didn't Start the Fire: Examining
the Stack Exchange Moderator and Contributor Strike*By *Yiwei Wu
(University of Texas at Austin)*Online communities and their host platforms
are mutually dependent yet conflict-prone. When platform policies clash
with community values, communities have resisted through strikes,
blackouts, and even migration to other platforms. Through such collective
actions, communities have sometimes won concessions, but these have
frequently proved to be temporary. Although previous research has
investigated strike events and migration chains, the processes by which
community-platform conflict unfolds remain obscure. How do
community-platform relationships deteriorate? How do communities organize
collective action? How do the participants proceed in the aftermath? We
investigate a conflict between the Stack Exchange platform and community
that occurred in 2023 around an emergency arising from the release of large
language models (LLMs). Based on a qualitative thematic analysis of 2,070
messages from Meta Stack Exchange and 14 interviews with community members,
we reveal how the 2023 conflict was preceded by a long-term deterioration
in the community-platform relationship, driven in particular by the
platform's disregard for the community's highly valued participatory role
in governance. Moreover, the platform's policy response to LLMs aggravated
the community's sense of crisis, triggering strike mobilization. We analyze
how the mobilization was coordinated through a tiered leadership and
communication structure, as well as how community members pivoted in the
aftermath. Building on recent theoretical scholarship in social computing,
we use Hirschman's exit, voice, and loyalty framework to theorize the
challenges of community-platform relations evinced in our data. Finally, we
recommend ways that platforms and communities can institute participatory
governance to be durable and effective.
Looking forward to seeing many of you,
Kinneret
--
Kinneret Gordon
Lead Research Community Officer
Wikimedia Foundation <https://wikimediafoundation.org/>
*Learn more about Wikimedia Research <https://research.wikimedia.org/>*
Hello,
Happy second week of February. This email describes two job postings that
may interest you or your colleagues. Earlier emails were sent to
Wiki-research-l
<https://lists.wikimedia.org/postorius/lists/wiki-research-l.lists.wikimedia…>
regarding these positions; please disregard the this email if you already
saw the info via Research-l. I'm consolidating info from both posts.
Position 1, posted at
https://www.demogr.mpg.de/en/career_6122/jobs_fellowships_1910/postdoctoral…,
is for a *postdoctoral researcher*. Quoting from the posting:
“Applications are invited from candidates *who have, or will soon obtain*,
a PhD in demography, sociology, epidemiology, medicine, health economics,
biostatistics, public health, or a related field.
"The successful candidate is expected to work within one or more of the
following areas:
1.
*Disease Presence and Disease Impact*
1.
Has disease presence become more or less predictive of disease
impact?
2.
How do the timing and patterns of disease accumulation shape the
onset of disease impact and individual health trajectories, including
pathways to death?
2.
*Diffusion of Medical Progress in Populations*
1.
How does medical progress shape health inequalities within
populations?
2.
How are changes in population health linked to how medical progress
is distributed within and across populations, for example through in- and
outpatient care?
3.
*Population Resilience and Vulnerability*
1.
How can population resilience and vulnerability be measured?
2.
Does medical progress make populations more or less vulnerable?”
Position 2 is for a *principal research scientist*. The job posting is at
https://job-boards.greenhouse.io/wikimedia/jobs/7597104. Quoting Leila (who
is CC'd on this email): “Please note
that if you have applied for the research scientist position which was
shared
<
https://lists.wikimedia.org/hyperkitty/list/wiki-research-l@lists.wikimedia…>
in December 2025, you are automatically considered for this new position
and there is no need to apply again. We will reach out to you if we need
additional information from you.”
Quoting from the job description:
“We are hiring a Principal Research Scientist to join the Wikimedia
Foundation’s Research team <https://research.wikimedia.org/team.html> to
support the Wikimedia communities and the Wikimedia Foundation in the
continued evolution of Wikimedia projects and their decentralized
governance model, ensuring that the projects become multigenerational
<https://meta.wikimedia.org/wiki/Strategy/Multigenerational>.
"Here are some things we’ve worked on that might give you a better sense of
the scope of the work you will be accountable for and part of:
-
Understanding Wikipedia administrator recruitment, retention, and
attrition (Learn more
<https://meta.wikimedia.org/wiki/Research:Wikipedia_Administrator_Recruitmen…>)
-
A set of recommendations for conducting NPOV research on Wikipedia (Learn
more
<https://meta.wikimedia.org/wiki/Research:Guidance_for_NPOV_Research_on_Wiki…>)
-
Developing a meta-method for analyzing the state of NPOV on Wikipedia (Learn
more
<https://meta.wikimedia.org/wiki/Research:A_meta-method_for_analyzing_NPOV_o…>)
"You can learn more about what the team has done in the past six months by
reading our biannual report <https://research.wikimedia.org/report.html>.
"Please note: this is a fully remote role that requires working with senior
leadership and stakeholders across the organization and the research
community in different timezones. It is expected that the Principal
Research Scientist be available for critical meetings and synchronous work
between 15:00 to 19:00 UTC. It is further expected that the candidate is
open to traveling up to four times per year.”
Leila posted some additional information on the Research mailing list. The
thread is viewable at
https://lists.wikimedia.org/hyperkitty/list/wiki-research-l@lists.wikimedia…
<https://lists.wikimedia.org/hyperkitty/list/wiki-research-l@lists.wikimedia…>
Good luck to any applicants.
(Disclaimer: I'm not an employee of either of the hiring organizations.
Please direct any questions to the applicable organization.)
Regards,
Pine🌲
Hi everyone,
The September 2025 Research Showcase will be live-streamed next Wednesday,
September 24, at 9:30 AM PT / 16:30 UTC. Find your local time here
<https://zonestamp.toolforge.org/1758731400>. Our theme this month is *Readers
and Readership Research*.
*We invite you to watch via the YouTube
stream: https://www.youtube.com/live/vqIp3CSgAVA
<https://www.youtube.com/live/vqIp3CSgAVA>.* As always, you can join the
conversation in the YouTube chat as soon as the showcase goes live.
This month, we will have one presentation featuring several speakers:
Key learnings from a decade of readers research and what’s ahead
By *WMF Research Team*
For over a decade, the Wikimedia Foundation's Research team and its
collaborators have studied Wikipedia's readership and readers. In this
showcase, we will synthesize the key insights from this body of work,
illustrate how these learnings inform our current research, and present
recommendations for future research into understanding Wikipedia's readers
and readership globally.
Looking forward to seeing many of you,
Kinneret
--
Kinneret Gordon
Lead Research Community Officer
Wikimedia Foundation <https://wikimediafoundation.org/>
*Learn more about Wikimedia Research <https://research.wikimedia.org/>*
Hi everyone,
The June 2025 Research Showcase will be live-streamed next Wednesday, June
18, at 9:30 AM PT / 16:30 UTC. Find your local time here
<https://zonestamp.toolforge.org/1750264200>. Our theme this month is *Ensuring
Content Integrity on Wikipedia*.
*We invite you to watch via the YouTube
stream: https://www.youtube.com/live/GgYh6zbrrss
<https://www.youtube.com/live/GgYh6zbrrss>.* As always, you can join the
conversation in the YouTube chat as soon as the showcase goes live.
Our presentations this month:
The Differential Effects of Page Protection on Wikipedia Article QualityBy
*Manoel Horta Ribeiro (Princeton University)*Wikipedia strives to be an
open platform where anyone can contribute, but that openness can sometimes
lead to conflicts or coordinated attempts to undermine article quality. To
address this, administrators use “page protection"—a tool that restricts
who can edit certain pages. But does this help the encyclopedia, or does it
do more harm than good? In this talk, I’ll present findings from a
large-scale, quasi-experimental study using over a decade of English
Wikipedia data. We focus on situations where editors requested page
protection and compare the outcomes for articles that were protected versus
similar ones that weren’t. Our results show that page protection has mixed
effects: it tends to benefit high-quality articles by preventing decline,
but it can hinder improvement in lower-quality ones. These insights reveal
how protection shapes Wikipedia content and help inform when it’s most
appropriate to restrict editing, and when it might be better to leave the
page open.
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
By
*Joshua Ashkinaze (University of Michigan)*Large language models (LLMs) are
trained on broad corpora and then used in communities with specialized
norms. Is providing LLMs with community rules enough for models to follow
these norms? We evaluate LLMs' capacity to detect (Task 1) and correct
(Task 2) biased Wikipedia edits according to Wikipedia's Neutral Point of
View (NPOV) policy. LLMs struggled with bias detection, achieving only 64%
accuracy on a balanced dataset. Models exhibited contrasting biases (some
under- and others over-predicted bias), suggesting distinct priors about
neutrality. LLMs performed better at generation, removing 79% of words
removed by Wikipedia editors. However, LLMs made additional changes beyond
Wikipedia editors' simpler neutralizations, resulting in high-recall but
low-precision editing. Interestingly, crowdworkers rated AI rewrites as
more neutral (70%) and fluent (61%) than Wikipedia-editor rewrites.
Qualitative analysis found LLMs sometimes applied NPOV more comprehensively
than Wikipedia editors but often made extraneous non-NPOV-related changes
(such as grammar). LLMs may apply rules in ways that resonate with the
public but diverge from community experts. While potentially effective for
generation, LLMs may reduce editor agency and increase moderation workload
(e.g., verifying additions). Even when rules are easy to articulate, having
LLMs apply them like community members may still be difficult.
Best,
Kinneret
--
Kinneret Gordon
Lead Research Community Officer
Wikimedia Foundation <https://wikimediafoundation.org/>
*Learn more about Wikimedia Research <https://research.wikimedia.org/>*