On Tue, Jul 7, 2026 at 4:26 PM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
And, surprise surprise, those [Google AI] summaries do cost clicks. See e.g. this Pew study from a year ago:
https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-lik...
Google and other search engines have been continuously optimizing those summaries since, so I would expect the effect size to increase over time. I personally start many of my searches on g.ai (their more robust AI Mode) directly; a lot of what we used to call "search" is IMO headed that way.
I was talking with a friend in the news business recently, and they said news sites are actively preparing for “Google Zero”, the day when Google Search sends them zero (or effectively zero, for revenue purposes) traffic. That’s why every quality news site is pushing subscriptions and social network video hard.
“Google Zero” is a very, very predictable future for us as well.[1]
A few months ago, before I had this useful and simple catchphrase, I urged some Foundation folks to orient their annual strategy process around the loss of 100% of search referral traffic, not 8% (as it was in the spring), or 20% (as it may be now), or whatever number it happens to be by the time we gather in Paris or Wikimania 2027.
But I think now that was the wrong question. The right question is: what are *we* are going to do about Google Zero?
This question is very explicitly not: *what is WMF going to do about Google Zero*. We can, should, must be more than WMF in this moment. If we’re reduced to just the capabilities and capacity of a single US non-profit of a few hundred people, we’re going to lose.
To win (or even just to survive) we must be a very broad, innovative, not-afraid-to-fail movement that happens to contain a US 501(c)3.
The good news, I think, is that we’re uniquely well-positioned to do that: - Open knowledge, aside from Wikipedia, has never been healthier (with arxiv, GH, etc.); - our APIs and data allow many experiments, as do those of our partners like Archive; - if we grow our editors through different channels, our mission can still be achieved even if we are no longer a top-10 website; - etc.
I hope to see a lot of you around Wikimania to talk about this.
Luis
[1] Not necessarily the only future; we’re currently on the path of newspapers (near-total collapse) but it isn’t impossible to imagine curves like arxiv’s or reddit’s or github’s https://lu.is/2026/04/web-collaboration-five-graphs/
It seems to me that we have to make a few uncomfortable choices before we start solving this.
Do we accept that our audience will shrink permanently and there is no way around it? Then we could embrace the role of a top source for LLMs and "ground truth". We would stop obsessing about the trends and focus on how to maintain a healthy community and funding.
Or do we want to fight against shrinking audiences? Then, in my mind, we would need to be very aggressive and start using some tools and technologies that we have always been very sceptical of. This, I guess, would include things like content optimisation, an AI chat, major investment in versatile social media channels and various new apps that deliver content under our brands to specific audiences.
D
На ср, 8.07.2026 г. в 3:41 Luis Villa via Wikimedia-l < wikimedia-l@lists.wikimedia.org> написа:
On Tue, Jul 7, 2026 at 4:26 PM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
And, surprise surprise, those [Google AI] summaries do cost clicks. See e.g. this Pew study from a year ago:
https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-lik...
Google and other search engines have been continuously optimizing those summaries since, so I would expect the effect size to increase over time. I personally start many of my searches on g.ai (their more robust AI Mode) directly; a lot of what we used to call "search" is IMO headed that way.
I was talking with a friend in the news business recently, and they said news sites are actively preparing for “Google Zero”, the day when Google Search sends them zero (or effectively zero, for revenue purposes) traffic. That’s why every quality news site is pushing subscriptions and social network video hard.
“Google Zero” is a very, very predictable future for us as well.[1]
A few months ago, before I had this useful and simple catchphrase, I urged some Foundation folks to orient their annual strategy process around the loss of 100% of search referral traffic, not 8% (as it was in the spring), or 20% (as it may be now), or whatever number it happens to be by the time we gather in Paris or Wikimania 2027.
But I think now that was the wrong question. The right question is: what are *we* are going to do about Google Zero?
This question is very explicitly not: *what is WMF going to do about Google Zero*. We can, should, must be more than WMF in this moment. If we’re reduced to just the capabilities and capacity of a single US non-profit of a few hundred people, we’re going to lose.
To win (or even just to survive) we must be a very broad, innovative, not-afraid-to-fail movement that happens to contain a US 501(c)3.
The good news, I think, is that we’re uniquely well-positioned to do that:
- Open knowledge, aside from Wikipedia, has never been healthier (with
arxiv, GH, etc.);
- our APIs and data allow many experiments, as do those of our partners
like Archive;
- if we grow our editors through different channels, our mission can still
be achieved even if we are no longer a top-10 website;
- etc.
I hope to see a lot of you around Wikimania to talk about this.
Luis
[1] Not necessarily the only future; we’re currently on the path of newspapers (near-total collapse) but it isn’t impossible to imagine curves like arxiv’s or reddit’s or github’s https://lu.is/2026/04/web-collaboration-five-graphs/
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
TLDR: WMF has to make hard choices, but we as a community, of which WMF is only one part, do not have to make false choices - we can try to do a lot of things.
On Wed, Jul 8, 2026 at 1:41 AM Dimitar Zagorski < dimitar.parvanov.dimitrov@gmail.com> wrote:
It seems to me that we have to make a few uncomfortable choices before we start solving this.
Do we accept that our audience will shrink permanently and there is no way around it? Then we could embrace the role of a top source for LLMs and "ground truth". We would stop obsessing about the trends and focus on how to maintain a healthy community and funding.
Or do we want to fight against shrinking audiences? Then, in my mind, we would need to be very aggressive and start using some tools and technologies that we have always been very sceptical of. This, I guess, would include things like content optimisation, an AI chat, major investment in versatile social media channels and various new apps that deliver content under our brands to specific audiences.
WMF, with large but finite resources, reputational concerns, etc etc will need to make hard choices. Maybe not exactly these hard choices, but some sort of hard choices.
**We as a broad movement with many members do not have to make this choice. We, the movement, can do both.** We can work to become a trusted source of ground truth; and work with new sources of traffic (like TikTok or YT or GLAMs or arxiv or all of the above) to drive traffic. Most of the "choices" ahead of us can be handled this way - we're tens of thousands of people, dozens of peer movement organizations, and thousands of knowledge-oriented fellow travelers. When someone says "we can only do one or the other" we should, as much as possible, reject that false binary and instead seek to find more resources, run more experiments, do more.
This will be easy, because open means we, the community, can just do stuff! we, the community, don’t have to ask permission! We can experiment, build, talk to each other as peers, expanding what works and dropping what doesn't. And we now have tooling that lets our best and most creative move faster than ever before.
And at the same time it will be hard. Frankly, we have let that open/BOLD habit atrophy, so we will need changes in practice and attitude.
Changes will have to include: (generalizations ahead, there are counter-examples to everything I say below but they are exceptions):
- community members need to stop thinking “step 0 is get funding from WMF” for every single activity. We need to get back to our roots as a scrappy, grassroots movement, which means both creativity in getting funding (wikicheese!) and creativity in building without funding. (I am keenly aware that this has equity impacts, and we need to be as thoughtful as possible about that, but we can't let that stop us from reinvigorating volunteering.)
- WMF needs to actively tell potential partners “open knowledge is in an all-hands-on-deck moment—so don’t treat WMF as the sole player”. Too many potential funders and partners look at WMF and think “We fund Wikimedia by funding WMF”/“WMF is the only way to partner with Wikipedia”. (I don’t think this is WMF’s fault, to be clear; it’s a natural pattern for funders to fall into.) But WMF needs to move from passively letting that happen, and towards actively disabusing funders and partners of that habit. All those partners need to get more comfortable partnering with chapters, individuals, or even doing things themselves.
- We need a culture of BOLD experimentation. The choice you’re saying we need to make, Dimi, would take a year or more and never really be settled. (It’s taking months just for the WMF to grapple with this internally, much less the movement as a whole.) So we need to as much as possible do “all of the above”, and build back our muscles of sharing experiments and learnings - including failing without fear, and sharing what failed. (I sincerely think one of the best things someone could do right now is become the Molly of AI+wiki, doing steady, consistent news-gathering and synthesis on what is working and what isn’t.)
- Scale of impact is important, but it isn’t everything. Sometimes we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a few million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
I’m sure there are others, and like I said, there are plenty of counterexamples to each of these. But they’re big patterns I see after a decade on the sidelines, and we need to unlearn them.
Hear hear! I warmly endorse Luis's last message. It matches my own observations, which I am expressing in my volunteer capacity, but basing on my experience as longtime staff as well.
I have been encouraging communities to develop and nurture a culture of experimentation https://commons.wikimedia.org/wiki/File:A_Culture_of_Experimentation_for_Wikimedia.pdf for several years now, and we have better tooling and infrastructure for experimentation than ever before.
I think one possible next step could be to identify a handful of experienced folks willing to be facilitators-and-mentors for groups of volunteers passionate about experimentation, then setting up such groups. The groups, with light facilitation, could generate ideas and testable hypotheses, document them, go ahead and experiment, and report back. (If this sounds appealing to folks, I would be willing to step up as one such facilitator/mentor, in my volunteer capacity.)
A.
On Wed, Jul 8, 2026 at 10:53 PM Luis Villa via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
TLDR: WMF has to make hard choices, but we as a community, of which WMF is only one part, do not have to make false choices - we can try to do a lot of things.
On Wed, Jul 8, 2026 at 1:41 AM Dimitar Zagorski < dimitar.parvanov.dimitrov@gmail.com> wrote:
It seems to me that we have to make a few uncomfortable choices before we start solving this.
Do we accept that our audience will shrink permanently and there is no way around it? Then we could embrace the role of a top source for LLMs and "ground truth". We would stop obsessing about the trends and focus on how to maintain a healthy community and funding.
Or do we want to fight against shrinking audiences? Then, in my mind, we would need to be very aggressive and start using some tools and technologies that we have always been very sceptical of. This, I guess, would include things like content optimisation, an AI chat, major investment in versatile social media channels and various new apps that deliver content under our brands to specific audiences.
WMF, with large but finite resources, reputational concerns, etc etc will need to make hard choices. Maybe not exactly these hard choices, but some sort of hard choices.
**We as a broad movement with many members do not have to make this choice. We, the movement, can do both.** We can work to become a trusted source of ground truth; and work with new sources of traffic (like TikTok or YT or GLAMs or arxiv or all of the above) to drive traffic. Most of the "choices" ahead of us can be handled this way - we're tens of thousands of people, dozens of peer movement organizations, and thousands of knowledge-oriented fellow travelers. When someone says "we can only do one or the other" we should, as much as possible, reject that false binary and instead seek to find more resources, run more experiments, do more.
This will be easy, because open means we, the community, can just do stuff! we, the community, don’t have to ask permission! We can experiment, build, talk to each other as peers, expanding what works and dropping what doesn't. And we now have tooling that lets our best and most creative move faster than ever before.
And at the same time it will be hard. Frankly, we have let that open/BOLD habit atrophy, so we will need changes in practice and attitude.
Changes will have to include: (generalizations ahead, there are counter-examples to everything I say below but they are exceptions):
- community members need to stop thinking “step 0 is get funding from WMF”
for every single activity. We need to get back to our roots as a scrappy, grassroots movement, which means both creativity in getting funding (wikicheese!) and creativity in building without funding. (I am keenly aware that this has equity impacts, and we need to be as thoughtful as possible about that, but we can't let that stop us from reinvigorating volunteering.)
- WMF needs to actively tell potential partners “open knowledge is in an
all-hands-on-deck moment—so don’t treat WMF as the sole player”. Too many potential funders and partners look at WMF and think “We fund Wikimedia by funding WMF”/“WMF is the only way to partner with Wikipedia”. (I don’t think this is WMF’s fault, to be clear; it’s a natural pattern for funders to fall into.) But WMF needs to move from passively letting that happen, and towards actively disabusing funders and partners of that habit. All those partners need to get more comfortable partnering with chapters, individuals, or even doing things themselves.
- We need a culture of BOLD experimentation. The choice you’re saying we
need to make, Dimi, would take a year or more and never really be settled. (It’s taking months just for the WMF to grapple with this internally, much less the movement as a whole.) So we need to as much as possible do “all of the above”, and build back our muscles of sharing experiments and learnings
- including failing without fear, and sharing what failed. (I sincerely
think one of the best things someone could do right now is become the Molly of AI+wiki, doing steady, consistent news-gathering and synthesis on what is working and what isn’t.)
- Scale of impact is important, but it isn’t everything. Sometimes we’ll
have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a few million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
I’m sure there are others, and like I said, there are plenty of counterexamples to each of these. But they’re big patterns I see after a decade on the sidelines, and we need to unlearn them.
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a few million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a few million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at all
the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping - https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a few million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)Based on stats total pageviews of human users of all wikis are dropping -https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7... https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7Cbar%7Call%7C(agent)~user%7Cmonthly
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Hi y'all, First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already). Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently. Cheers, Nicolas Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l <wikimedia-l@lists.wikimedia.org> a écrit : On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote: > - Scale of impact is important, but it isn’t everything. Sometimes > we’ll have to build things without knowing if they’ll scale, or even > perhaps knowing that they won’t scale (eg Depths of Wikipedia is one > of the best community-builders we have even though it has “only” a few > million followers). “That’s good but it won’t scale to enwiki” can > never be allowed to be a blocker for anyone. If the experiment is > actually great, then we’ll figure out a way to get it onto enwiki. Or > not! That shouldn’t stop people. At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed. --Michael Snow _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/message/6EQOFSG34E3BC6FIYL55HYBTQBAW4XFN/ To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/message/4WPCSTZNDMBDMYHWHV27BN2EZLGY5LXB/ To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list --wikimedia-l@lists.wikimedia.org, guidelines at:https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines andhttps://meta.wikimedia.org/wiki/Wikimedia-l Public archives athttps://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email towikimedia-l-leave@lists.wikimedia.org
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at all
the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a few million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
" it looks like Wikipedia will be mainly the base for training AI" I would like to challenge that claim. It is less and less true by the day and if we keep on losing editors it also means our content will be less up to date, less complete, less multilingual, less localized which all lead to less interesting training datasets.
OpenAI has already started to invest in local journalism because they need fresh, actual and local content to make sure their model improves.
Which gives hope, we're years ahead from them but we're not on a good trend on that either.
-- Christophe
On Thu, 9 Jul 2026 at 10:18, Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at
all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a
few
million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at
all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a
few
million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki. Or not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
I realize the intent of Google Zero appears like a cynical action keeping everyone within the Google world, which will provide Google with strong revenue streams. I think the reality is that many people already dont trust the Googlesphere and those people will move away from it. That portion must be our major target audience, while hopefully the Enterprise project will be able to increase revenue from the movement Google for faster more reliable access.
As Google goes down this path we can look forward to measures that break Googles monolithic setup through legal channels in many countries especially in Europe, and those of competing organisations. Its not all doom and gloom. It will need a more concerted effort with a lot of outside the box thinking that can enhance the Wikimedia movement's position as primary trusted knowledge source.
On Thu, 9 Jul 2026 at 16:39, James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at
all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a
few
million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki.
Or
not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
James, thank you for grounding this with real numbers rather than impressions — the Similarweb figure and the cross-language health content analysis are exactly the kind of evidence this thread needs, and they're worth taking as the starting point.
That said, I don't think there's much left to gain from debating what should have been done differently. What happened, happened — the past doesn't come back, and no amount of hindsight changes the trajectory from here.
What matters more now is the immediate ahead of us, and looking at it with a different mindset: not only as a problem, but as something that might open an opportunity — because very often what looks like a threat to us is somebody else's opportunity, and figuring out which side of that split we're on matters more than assigning blame for the past.
A few things, connected, that point in that direction (even not a complete picture). AI infrastructure spending is running above $1 trillion globally this year, with investors increasingly nervous about the return on it. [1] At the same time, open-weight models that run locally — largely out of China, but not only — are now within a few points of frontier benchmarks at a fraction of the cost.[2] Regulation is starting to require actual, checkable data provenance rather than declared compliance (EU AI Act for instance). And large parts of the web are moving toward walling off AI crawlers and charging for metered access by default.[3]
Put together, I think James's numbers might not just describe a decline — they might be the leading edge of a market starting to separate "content that's cheap to grab" from "content whose origin can be verified," and pricing them very differently. If that's right, it matters most exactly where James's data points: health content is where the cost of being wrong is highest, and where verifiable sourcing should be worth the most, not the least.
kind regards
Ilario [1] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-a... [2] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://computingforgeeks.com/open-source-llm-comparison/ [3 https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies... https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/
On Thu, Jul 9, 2026 at 10:39 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at
all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote:
- Scale of impact is important, but it isn’t everything. Sometimes
we’ll have to build things without knowing if they’ll scale, or even perhaps knowing that they won’t scale (eg Depths of Wikipedia is one of the best community-builders we have even though it has “only” a
few
million followers). “That’s good but it won’t scale to enwiki” can never be allowed to be a blocker for anyone. If the experiment is actually great, then we’ll figure out a way to get it onto enwiki.
Or
not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
EU data provenance regulations do not require foundation model providers to link to the constituent sources of AI model knowledge. In fact the disclosure rules don’t even apply to foundation models themselves yet, only services. And they only require disclosing that text or other media were generated by AI (via the interface and metadata). It’s not going to help fix brand awareness and loss in referral traffic, only transparency that a particular piece of content is or is not AI generated.
Google already links to Wikipedia when it is used as a source in search answers, and we’re still seeing a traffic decline. We shouldn’t wait to assume that EU regulations or walled gardens are going to save us from being disintermediated.
The drop in enwiki traffic in particular is likely indicative of a structural shift considering that in just two years (2024-2026) we have gone from 33% of Americans using chatbots and AI summaries to 49% ( https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbo... )
Steven Walling
On Thu, Jul 9, 2026 at 7:39 AM Ilario Valdelli via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
James, thank you for grounding this with real numbers rather than impressions — the Similarweb figure and the cross-language health content analysis are exactly the kind of evidence this thread needs, and they're worth taking as the starting point.
That said, I don't think there's much left to gain from debating what should have been done differently. What happened, happened — the past doesn't come back, and no amount of hindsight changes the trajectory from here.
What matters more now is the immediate ahead of us, and looking at it with a different mindset: not only as a problem, but as something that might open an opportunity — because very often what looks like a threat to us is somebody else's opportunity, and figuring out which side of that split we're on matters more than assigning blame for the past.
A few things, connected, that point in that direction (even not a complete picture). AI infrastructure spending is running above $1 trillion globally this year, with investors increasingly nervous about the return on it. [1] At the same time, open-weight models that run locally — largely out of China, but not only — are now within a few points of frontier benchmarks at a fraction of the cost.[2] Regulation is starting to require actual, checkable data provenance rather than declared compliance (EU AI Act for instance). And large parts of the web are moving toward walling off AI crawlers and charging for metered access by default.[3]
Put together, I think James's numbers might not just describe a decline — they might be the leading edge of a market starting to separate "content that's cheap to grab" from "content whose origin can be verified," and pricing them very differently. If that's right, it matters most exactly where James's data points: health content is where the cost of being wrong is highest, and where verifiable sourcing should be worth the most, not the least.
kind regards
Ilario [1] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-a... [2] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://computingforgeeks.com/open-source-llm-comparison/ [3 https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies...
On Thu, Jul 9, 2026 at 10:39 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at
all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote: > - Scale of impact is important, but it isn’t everything. Sometimes > we’ll have to build things without knowing if they’ll scale, or even > perhaps knowing that they won’t scale (eg Depths of Wikipedia is one > of the best community-builders we have even though it has “only” a few > million followers). “That’s good but it won’t scale to enwiki” can > never be allowed to be a blocker for anyone. If the experiment is > actually great, then we’ll figure out a way to get it onto enwiki. Or > not! That shouldn’t stop people.
At this point, no single initiative (whether community, chapter, or WMF-led) will have an inherent capacity to scale to enwiki. To do so would require corporate-style resources with a corresponding top-down mandate for implementation, and risk (rightly) triggering resistance to the point of launching a true durable fork, for once. Instead, the best approach would be to find ideas that are good enough to propagate, encourage them, and when they have propagated enough, they will scale with relatively little intervention needed.
--Michael Snow
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- Ilario Valdelli Skype: valdelli Tel: +41764821371 http://www.wikimedia.ch _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hi,
One note that hasn't been mentioned is that the future UI for consuming information will be "agentic," where the tool used will search for and collect the latest information on the topics. This means systems will prioritize content where the information source is described (i.e., references), as accuracy can be validated from there. So references are one of the strengths of Wikimedia services compared to almost all other sites on the internet. This doesn't really solve the eyeball problem since the content is still used externally, but it should be good for something.
Br, -- Kimmo Virtanen, Zache
On Fri, Jul 10, 2026 at 3:14 AM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
EU data provenance regulations do not require foundation model providers to link to the constituent sources of AI model knowledge. In fact the disclosure rules don’t even apply to foundation models themselves yet, only services. And they only require disclosing that text or other media were generated by AI (via the interface and metadata). It’s not going to help fix brand awareness and loss in referral traffic, only transparency that a particular piece of content is or is not AI generated.
Google already links to Wikipedia when it is used as a source in search answers, and we’re still seeing a traffic decline. We shouldn’t wait to assume that EU regulations or walled gardens are going to save us from being disintermediated.
The drop in enwiki traffic in particular is likely indicative of a structural shift considering that in just two years (2024-2026) we have gone from 33% of Americans using chatbots and AI summaries to 49% ( https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbo... )
Steven Walling
On Thu, Jul 9, 2026 at 7:39 AM Ilario Valdelli via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
James, thank you for grounding this with real numbers rather than impressions — the Similarweb figure and the cross-language health content analysis are exactly the kind of evidence this thread needs, and they're worth taking as the starting point.
That said, I don't think there's much left to gain from debating what should have been done differently. What happened, happened — the past doesn't come back, and no amount of hindsight changes the trajectory from here.
What matters more now is the immediate ahead of us, and looking at it with a different mindset: not only as a problem, but as something that might open an opportunity — because very often what looks like a threat to us is somebody else's opportunity, and figuring out which side of that split we're on matters more than assigning blame for the past.
A few things, connected, that point in that direction (even not a complete picture). AI infrastructure spending is running above $1 trillion globally this year, with investors increasingly nervous about the return on it. [1] At the same time, open-weight models that run locally — largely out of China, but not only — are now within a few points of frontier benchmarks at a fraction of the cost.[2] Regulation is starting to require actual, checkable data provenance rather than declared compliance (EU AI Act for instance). And large parts of the web are moving toward walling off AI crawlers and charging for metered access by default.[3]
Put together, I think James's numbers might not just describe a decline — they might be the leading edge of a market starting to separate "content that's cheap to grab" from "content whose origin can be verified," and pricing them very differently. If that's right, it matters most exactly where James's data points: health content is where the cost of being wrong is highest, and where verifiable sourcing should be worth the most, not the least.
kind regards
Ilario [1] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-a... [2] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://computingforgeeks.com/open-source-llm-comparison/ [3 https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies...
On Thu, Jul 9, 2026 at 10:39 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look at
all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi y'all,
First could someone remind me how much traffic comes from Google (I have a vague memory from last year that this is already very low - maybe 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if true we essentially are at Google Zero already).
Secondly, I'd love to see initiative outside of enwiki. If you look at all the ~1000 Wikimedia projects, almost all of them are seeing an increase in traffic ; only a few big old wikipedias are seeing a decrease (which are also the ones with the most traffic, hence why it's more visible but it obfuscate at the big picture). I think we should start by looking at and learning from the "successful" projects instead of focusing on the declining ones. For instance, Wikisources always have been badly to not-at-all referenced by Google and yet there is a strong increase in pageviews recently.
Cheers, Nicolas
Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> a écrit :
> On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote: > > - Scale of impact is important, but it isn’t everything. Sometimes > > we’ll have to build things without knowing if they’ll scale, or > even > > perhaps knowing that they won’t scale (eg Depths of Wikipedia is > one > > of the best community-builders we have even though it has “only” a > few > > million followers). “That’s good but it won’t scale to enwiki” can > > never be allowed to be a blocker for anyone. If the experiment is > > actually great, then we’ll figure out a way to get it onto enwiki. > Or > > not! That shouldn’t stop people. > > At this point, no single initiative (whether community, chapter, or > WMF-led) will have an inherent capacity to scale to enwiki. To do so > would require corporate-style resources with a corresponding > top-down > mandate for implementation, and risk (rightly) triggering resistance > to > the point of launching a true durable fork, for once. Instead, the > best > approach would be to find ideas that are good enough to propagate, > encourage them, and when they have propagated enough, they will > scale > with relatively little intervention needed. > > --Michael Snow > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- Ilario Valdelli Skype: valdelli Tel: +41764821371 http://www.wikimedia.ch _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hoi. Given the lack of maintenance of many Wikipedia articles, eg the median time for a correction of a retracted reference is 3.68 years, the notion that our references will save us is unlikely. Given how big we are on references, we could expose them in our own search engine where we combine what we have, linked to subjects and enable people to add sources and work on our content..
This is not a Wikipedia centred idea and I expect that it will not be considered. But hey, what other ideas are out there that have a chance to expose people to our content? Thanks, GerardM
On Fri, 10 Jul 2026 at 03:26, Kimmo Virtanen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi,
One note that hasn't been mentioned is that the future UI for consuming information will be "agentic," where the tool used will search for and collect the latest information on the topics. This means systems will prioritize content where the information source is described (i.e., references), as accuracy can be validated from there. So references are one of the strengths of Wikimedia services compared to almost all other sites on the internet. This doesn't really solve the eyeball problem since the content is still used externally, but it should be good for something.
Br, -- Kimmo Virtanen, Zache
On Fri, Jul 10, 2026 at 3:14 AM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
EU data provenance regulations do not require foundation model providers to link to the constituent sources of AI model knowledge. In fact the disclosure rules don’t even apply to foundation models themselves yet, only services. And they only require disclosing that text or other media were generated by AI (via the interface and metadata). It’s not going to help fix brand awareness and loss in referral traffic, only transparency that a particular piece of content is or is not AI generated.
Google already links to Wikipedia when it is used as a source in search answers, and we’re still seeing a traffic decline. We shouldn’t wait to assume that EU regulations or walled gardens are going to save us from being disintermediated.
The drop in enwiki traffic in particular is likely indicative of a structural shift considering that in just two years (2024-2026) we have gone from 33% of Americans using chatbots and AI summaries to 49% ( https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbo... )
Steven Walling
On Thu, Jul 9, 2026 at 7:39 AM Ilario Valdelli via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
James, thank you for grounding this with real numbers rather than impressions — the Similarweb figure and the cross-language health content analysis are exactly the kind of evidence this thread needs, and they're worth taking as the starting point.
That said, I don't think there's much left to gain from debating what should have been done differently. What happened, happened — the past doesn't come back, and no amount of hindsight changes the trajectory from here.
What matters more now is the immediate ahead of us, and looking at it with a different mindset: not only as a problem, but as something that might open an opportunity — because very often what looks like a threat to us is somebody else's opportunity, and figuring out which side of that split we're on matters more than assigning blame for the past.
A few things, connected, that point in that direction (even not a complete picture). AI infrastructure spending is running above $1 trillion globally this year, with investors increasingly nervous about the return on it. [1] At the same time, open-weight models that run locally — largely out of China, but not only — are now within a few points of frontier benchmarks at a fraction of the cost.[2] Regulation is starting to require actual, checkable data provenance rather than declared compliance (EU AI Act for instance). And large parts of the web are moving toward walling off AI crawlers and charging for metered access by default.[3]
Put together, I think James's numbers might not just describe a decline — they might be the leading edge of a market starting to separate "content that's cheap to grab" from "content whose origin can be verified," and pricing them very differently. If that's right, it matters most exactly where James's data points: health content is where the cost of being wrong is highest, and where verifiable sourcing should be worth the most, not the least.
kind regards
Ilario [1] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-a... [2] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://computingforgeeks.com/open-source-llm-comparison/ [3 https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies...
On Thu, Jul 9, 2026 at 10:39 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I agree we ought to analyze better our situation, before we run to conclusions and discuss remedies
Swwp has seen a decrease in traffic like others but in June we have seen an increase even if we include only accesses from mobiles (at least for June)
https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm...
What we have seen more clearly though is that accesses of the most read articles per day in decreasing.
We have also more subjectively from media, and social media seen a growing skepticism toward AI-generated results and an increase in the trust of Wikipedia results
Our general conclusions this far are
*We are doing well in our core encyclopedic articles (like of administrative divisions within a country) , but less well in newslike articles (which the internet readers asks more frequently, but where newswebsites always been stronger). And we should concentrate our efforts where we are strong
*We are in a late stage in our general product life cycle and must accept this. We will never have the increase in editors etc like we hade 15 yeas ago, but must do what we can do thrive in a this mode. For us - less angry internal argument, partially done by avoiding articles of infected subjects like Israel-Palestine
And for Google Zero. I believe it is them who has a problem not us. Early on, and still, all new PC user were forced into Microsoft Edge and Bing which lead us to Firefox and Google. Google zero as I see will lead to an exodus från Google, there are many other search engines that works fine
Wikipedia is great
Anders
Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l:
Hi,
Secondly, I'd love to see initiative outside of enwiki. If you look > at all the ~1000 Wikimedia projects, almost all of them are seeing an > increase in traffic ; only a few big old wikipedias are seeing a decrease > (which are also the ones with the most traffic, hence why it's more visible > but it obfuscate at the big picture)
Based on stats total pageviews of human users of all wikis are dropping
https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7...
Br, -- Kimmo Virtanen, Zache
On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> Hi y'all, > > First could someone remind me how much traffic comes from Google (I > have a vague memory from last year that this is already very low - maybe 10 > or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if > true we essentially are at Google Zero already). > > Secondly, I'd love to see initiative outside of enwiki. If you look > at all the ~1000 Wikimedia projects, almost all of them are seeing an > increase in traffic ; only a few big old wikipedias are seeing a decrease > (which are also the ones with the most traffic, hence why it's more visible > but it obfuscate at the big picture). > I think we should start by looking at and learning from the > "successful" projects instead of focusing on the declining ones. For > instance, Wikisources always have been badly to not-at-all referenced by > Google and yet there is a strong increase in pageviews recently. > > Cheers, > Nicolas > > Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < > wikimedia-l@lists.wikimedia.org> a écrit : > >> On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote: >> > - Scale of impact is important, but it isn’t everything. >> Sometimes >> > we’ll have to build things without knowing if they’ll scale, or >> even >> > perhaps knowing that they won’t scale (eg Depths of Wikipedia is >> one >> > of the best community-builders we have even though it has “only” >> a few >> > million followers). “That’s good but it won’t scale to enwiki” >> can >> > never be allowed to be a blocker for anyone. If the experiment is >> > actually great, then we’ll figure out a way to get it onto >> enwiki. Or >> > not! That shouldn’t stop people. >> >> At this point, no single initiative (whether community, chapter, or >> WMF-led) will have an inherent capacity to scale to enwiki. To do >> so >> would require corporate-style resources with a corresponding >> top-down >> mandate for implementation, and risk (rightly) triggering >> resistance to >> the point of launching a true durable fork, for once. Instead, the >> best >> approach would be to find ideas that are good enough to propagate, >> encourage them, and when they have propagated enough, they will >> scale >> with relatively little intervention needed. >> >> --Michael Snow >> >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >> guidelines at: >> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to >> wikimedia-l-leave@lists.wikimedia.org > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- Ilario Valdelli Skype: valdelli Tel: +41764821371 http://www.wikimedia.ch _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Yah a search engine that actually gives real references that supports the statements in question would be amazing. Google currently gives mostly fake references (ie the references they provide often do not support the content it is attached behind).
J
On Fri, Jul 10, 2026 at 7:20 AM Gerard Meijssen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hoi. Given the lack of maintenance of many Wikipedia articles, eg the median time for a correction of a retracted reference is 3.68 years, the notion that our references will save us is unlikely. Given how big we are on references, we could expose them in our own search engine where we combine what we have, linked to subjects and enable people to add sources and work on our content..
This is not a Wikipedia centred idea and I expect that it will not be considered. But hey, what other ideas are out there that have a chance to expose people to our content? Thanks, GerardM
On Fri, 10 Jul 2026 at 03:26, Kimmo Virtanen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi,
One note that hasn't been mentioned is that the future UI for consuming information will be "agentic," where the tool used will search for and collect the latest information on the topics. This means systems will prioritize content where the information source is described (i.e., references), as accuracy can be validated from there. So references are one of the strengths of Wikimedia services compared to almost all other sites on the internet. This doesn't really solve the eyeball problem since the content is still used externally, but it should be good for something.
Br, -- Kimmo Virtanen, Zache
On Fri, Jul 10, 2026 at 3:14 AM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
EU data provenance regulations do not require foundation model providers to link to the constituent sources of AI model knowledge. In fact the disclosure rules don’t even apply to foundation models themselves yet, only services. And they only require disclosing that text or other media were generated by AI (via the interface and metadata). It’s not going to help fix brand awareness and loss in referral traffic, only transparency that a particular piece of content is or is not AI generated.
Google already links to Wikipedia when it is used as a source in search answers, and we’re still seeing a traffic decline. We shouldn’t wait to assume that EU regulations or walled gardens are going to save us from being disintermediated.
The drop in enwiki traffic in particular is likely indicative of a structural shift considering that in just two years (2024-2026) we have gone from 33% of Americans using chatbots and AI summaries to 49% ( https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbo... )
Steven Walling
On Thu, Jul 9, 2026 at 7:39 AM Ilario Valdelli via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
James, thank you for grounding this with real numbers rather than impressions — the Similarweb figure and the cross-language health content analysis are exactly the kind of evidence this thread needs, and they're worth taking as the starting point.
That said, I don't think there's much left to gain from debating what should have been done differently. What happened, happened — the past doesn't come back, and no amount of hindsight changes the trajectory from here.
What matters more now is the immediate ahead of us, and looking at it with a different mindset: not only as a problem, but as something that might open an opportunity — because very often what looks like a threat to us is somebody else's opportunity, and figuring out which side of that split we're on matters more than assigning blame for the past.
A few things, connected, that point in that direction (even not a complete picture). AI infrastructure spending is running above $1 trillion globally this year, with investors increasingly nervous about the return on it. [1] At the same time, open-weight models that run locally — largely out of China, but not only — are now within a few points of frontier benchmarks at a fraction of the cost.[2] Regulation is starting to require actual, checkable data provenance rather than declared compliance (EU AI Act for instance). And large parts of the web are moving toward walling off AI crawlers and charging for metered access by default.[3]
Put together, I think James's numbers might not just describe a decline — they might be the leading edge of a market starting to separate "content that's cheap to grab" from "content whose origin can be verified," and pricing them very differently. If that's right, it matters most exactly where James's data points: health content is where the cost of being wrong is highest, and where verifiable sourcing should be worth the most, not the least.
kind regards
Ilario [1] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-a... [2] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://computingforgeeks.com/open-source-llm-comparison/ [3 https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies...
On Thu, Jul 9, 2026 at 10:39 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
I do not think Google Zero will be (is) experienced by anyone except for us as a problem.
Sure, the current AI search is not perfect, but it will be much better on a horizon of months. It is unfortunate that the end users take the results uncritically rather than checking them clicking on the source links, but it is what it is. May be this culture would change, but this timescale is way way longer than the time to improve the AI search quality.
Except for obvious financial implications (which are very important though) it looks like Wikipedia will be mainly the base for training AI. It is ok, in real life most people do not read encyclopaedias. It is important to be the training base as well, but we should also realize that if this is our role then some priorities must be shifted. For example, AI do not care whether the article looks nice or not, they do not care about usability. They probably do not care about categories and many other things we spent years trying to make them perfect. They do care about content though.
On the other hand, there are clearly things AI can not do, which were already mentioned here. They can not read offline material, and much of the paywall material. They can not produce pictures out of nothing - which means Commons might have a very different role from what it has now. It might even become the flagship project (unlikely though with the current Commons community). But all these things require attention, and many of them will be resisted by the community in the first place.
Best Yaroslav
On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> I agree we ought to analyze better our situation, before we run to > conclusions and discuss remedies > > Swwp has seen a decrease in traffic like others but in June we have > seen an increase even if we include only accesses from mobiles (at least > for June) > > > https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm... > > What we have seen more clearly though is that accesses of the most > read articles per day in decreasing. > > We have also more subjectively from media, and social media seen a > growing skepticism toward AI-generated results and an increase in the trust > of Wikipedia results > > Our general conclusions this far are > > *We are doing well in our core encyclopedic articles (like of > administrative divisions within a country) , but less well in newslike > articles (which the internet readers asks more frequently, but where > newswebsites always been stronger). And we should concentrate our efforts > where we are strong > > *We are in a late stage in our general product life cycle and must > accept this. We will never have the increase in editors etc like we hade 15 > yeas ago, but must do what we can do thrive in a this mode. For us - less > angry internal argument, partially done by avoiding articles of infected > subjects like Israel-Palestine > > And for Google Zero. I believe it is them who has a problem not us. > Early on, and still, all new PC user were forced into Microsoft Edge and > Bing which lead us to Firefox and Google. Google zero as I see will lead > to an exodus från Google, there are many other search engines that works > fine > > Wikipedia is great > > Anders > > > > > Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l: > > Hi, > > Secondly, I'd love to see initiative outside of enwiki. If you look >> at all the ~1000 Wikimedia projects, almost all of them are seeing an >> increase in traffic ; only a few big old wikipedias are seeing a decrease >> (which are also the ones with the most traffic, hence why it's more visible >> but it obfuscate at the big picture) > > > Based on stats total pageviews of human users of all wikis are > dropping > - > https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7... > > > Br, > -- Kimmo Virtanen, Zache > > On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < > wikimedia-l@lists.wikimedia.org> wrote: > >> Hi y'all, >> >> First could someone remind me how much traffic comes from Google (I >> have a vague memory from last year that this is already very low - maybe 10 >> or 20 % - and that LLMs are already bringing more traffic to Wikipedia, if >> true we essentially are at Google Zero already). >> >> Secondly, I'd love to see initiative outside of enwiki. If you look >> at all the ~1000 Wikimedia projects, almost all of them are seeing an >> increase in traffic ; only a few big old wikipedias are seeing a decrease >> (which are also the ones with the most traffic, hence why it's more visible >> but it obfuscate at the big picture). >> I think we should start by looking at and learning from the >> "successful" projects instead of focusing on the declining ones. For >> instance, Wikisources always have been badly to not-at-all referenced by >> Google and yet there is a strong increase in pageviews recently. >> >> Cheers, >> Nicolas >> >> Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < >> wikimedia-l@lists.wikimedia.org> a écrit : >> >>> On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote: >>> > - Scale of impact is important, but it isn’t everything. >>> Sometimes >>> > we’ll have to build things without knowing if they’ll scale, or >>> even >>> > perhaps knowing that they won’t scale (eg Depths of Wikipedia is >>> one >>> > of the best community-builders we have even though it has “only” >>> a few >>> > million followers). “That’s good but it won’t scale to enwiki” >>> can >>> > never be allowed to be a blocker for anyone. If the experiment >>> is >>> > actually great, then we’ll figure out a way to get it onto >>> enwiki. Or >>> > not! That shouldn’t stop people. >>> >>> At this point, no single initiative (whether community, chapter, >>> or >>> WMF-led) will have an inherent capacity to scale to enwiki. To do >>> so >>> would require corporate-style resources with a corresponding >>> top-down >>> mandate for implementation, and risk (rightly) triggering >>> resistance to >>> the point of launching a true durable fork, for once. Instead, the >>> best >>> approach would be to find ideas that are good enough to propagate, >>> encourage them, and when they have propagated enough, they will >>> scale >>> with relatively little intervention needed. >>> >>> --Michael Snow >>> >>> _______________________________________________ >>> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >>> guidelines at: >>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>> https://meta.wikimedia.org/wiki/Wikimedia-l >>> Public archives at >>> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >>> To unsubscribe send an email to >>> wikimedia-l-leave@lists.wikimedia.org >> >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >> guidelines at: >> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to >> wikimedia-l-leave@lists.wikimedia.org > > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- Ilario Valdelli Skype: valdelli Tel: +41764821371 http://www.wikimedia.ch _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Sadly, the only realistic path I see there would be through acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik
Hoi, It depends what we aim to be. When we expose our references as a start, when we combine it with what we also have in media and data. When the media is presented, when the data is presented in an intelligible way, we already serve the sum of all the knowledge as we have it. We do not pretend to serve in competition to a Google, we serve enriched information and Wikipedia is very much served to a new public from a different angle.
When we finally REALLY partner with like minded orgs like Internet Archive, Open Library, the Biodiversity Heritage Library ... results to their websites will feature equally prominent. When we REALLY partner with ORCiD, CrossRef, Retraction Watch our long awaited WikiCite will finally happen.
The way to start is by starting in a Wiki way as we did with Commons. It started its physical existence like an idea, it then morphed into a working concept and finally into the project it is today. We do not have to be perfect, it will not be perfect but we sure as hell can be functional and draw its own crowd. A crowd that will work in an open environment and invited to make it different, different for the better and the crowd may become a more global community. Thanks, GerardM
On Fri, 10 Jul 2026 at 18:52, Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Sadly, the only realistic path I see there would be through acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hoi, How does that raise more eyeballs to our data, to our projects. Why would we build this to serve companies that raise CO2 with 35%?
Why support AI at all when it is killing our future! Thanks, GerardM
On Fri, 10 Jul 2026 at 19:30, Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
On Fri, Jul 10, 2026 at 10:36 AM Gerard Meijssen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hoi, How does that raise more eyeballs to our data, to our projects. Why would we build this to serve companies that raise CO2 with 35%?
Why support AI at all when it is killing our future! Thanks, GerardM
Gerard,
ChatGPT alone has over a billion monthly active users and these platforms are growing rapidly (note the survey data I linked to previously in thread).
Everything Wikimedia does has a CO2 emissions cost. We are not carbon neutral today as far as I am aware, though I would bet many of us would agree that is a goal worth pursuing.
Steven
On Fri, 10 Jul 2026 at 19:30, Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at
https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hoi, You indicate that over a billion people use ChatGPT... the closest analogy is the Boxer wars where the UK fought to have a market for opium in China. It was to keep people depending on drugs. AI from a strategic point of view creates a dependency on the USA. The global narrative is that as a partner, the USA can no longer be depended on. It is why Europe and China are moving away from products that are keyed to the US national interest including products like SWIFT, personal computing, computing.
When some US based megalomanic billionaires outcompete each other to develop AI, they are a substantial producer of worldwide CO2, this while we are already past tipping points we should not cross when we want to have a livable earth. When it is considered a fact that a billion people use ChatGPT, I wonder if that is because a lot of the hardware and software is infused with AI.
AI production is not only producing CO2 it is actively using up water where there is not enough to go around in the quantity desired by AI factories. There is a reason why people do not want AI factories in their backyard and there is a reason why these factories are announced when permissions have already been granted.
Google used to use green energy for its services, it no longer can.. the same is true for all AI companies.
AI is used in a strategic fashion and it cannot be assumed to be/remain globally available. We are a global organisation and our public is global. AI, in its current iteration, is something we cannot afford. When we provide a search engine with all our references, with inclusion of like minded organisations, when videos are included and when we enable annotation to all we offer, we do not offer a substitute for AI but we offer something substantially more inclusive than just Wikipedia.
Also, when the WMF wants to be green, it can source its energy from a supplier that provides green energy. When that is not possible, it can build an equivalent new capacity and provide it on the grid. Then again, in this argument we consider AI not WMF. Thanks, GerardM
On Fri, 10 Jul 2026 at 20:05, Steven Walling steven.walling@gmail.com wrote:
On Fri, Jul 10, 2026 at 10:36 AM Gerard Meijssen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hoi, How does that raise more eyeballs to our data, to our projects. Why would we build this to serve companies that raise CO2 with 35%?
Why support AI at all when it is killing our future! Thanks, GerardM
Gerard,
ChatGPT alone has over a billion monthly active users and these platforms are growing rapidly (note the survey data I linked to previously in thread).
Everything Wikimedia does has a CO2 emissions cost. We are not carbon neutral today as far as I am aware, though I would bet many of us would agree that is a goal worth pursuing.
Steven
On Fri, 10 Jul 2026 at 19:30, Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l
Public archives at
https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts (i.e. blog posts by authoritative companies or recently published ScienceDirect articles) - Content that has been updated recently, with the biggest "hot takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months) - Content that helps users make a decision between different choices (i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other authorities - We ground our content in anonymity instead of named experts or instutional process/opinoin - We rarely do original analysis instead summarizing the experience and expertise of others, - Alot of our content is out of date, and self-aware of its gaps (i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them: https://ahrefs.com/blog/most-cited-domains-ai-overviews/ - I have access to SERanking's corpus of prompt monitoring across 5 models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps. - Studies from earlier in the year put Wikipedia at about ~13% of CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well) - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance) - Domain specific citation pools have pretty significant differences in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones? - How many of the interactions are two or three steps down a chain of more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ? - How much are the AI companies optimizing for "sales" or "addiction" rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches? - How much is geolocation forcing more and more responses into "local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts
(i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot takes"
(i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different choices
(i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the experience
and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps (i.e.
maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across 5
models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend
to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant differences
in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain of
more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or "addiction"
rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into "local"
sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts
(i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot
takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different choices
(i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other
authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the experience
and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps
(i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across 5
models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend
to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant differences
in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain of
more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or "addiction"
rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into
"local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that supports
the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles
On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts
(i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot
takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different choices
(i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other
authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the experience
and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps
(i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across 5
models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend
to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant differences
in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain
of more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or
"addiction" rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into
"local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that
supports the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
English Wikipedia is a conservative organization / project as are many other large versions of Wikipedia. Many stakeholders need to be convinced / brought on board to make even relatively minor changes. Innovating within smaller versions of Wikipedia or in other projects outside Wikipedia is much easier. And successes can occasionally be brought into the larger Wikipedias such as we did with Our World in Data interactive graphs...
https://en.wikipedia.org/wiki/Wheat#Production_and_consumption
We have now added 100s of these in various languages. And they are getting thousands of plays a day.
James
On Sat, Jul 11, 2026 at 3:28 PM Charles Roberson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles
On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts
(i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot
takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different
choices (i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other
authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the experience
and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps
(i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across 5
models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn
tend to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant
differences in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain
of more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or
"addiction" rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into
"local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
> Yah a search engine that actually gives real references that supports the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Are you talking about that godawful thing that I just clicked on in the "wheat" article, that the back button doesn't work on?
I'm pulling that out. Sorry, but "back button works" is a basic thing of Web functionality. I should not need to click a "return to article" button to get back where I was.
Todd
On Sat, Jul 11, 2026 at 8:32 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
English Wikipedia is a conservative organization / project as are many other large versions of Wikipedia. Many stakeholders need to be convinced / brought on board to make even relatively minor changes. Innovating within smaller versions of Wikipedia or in other projects outside Wikipedia is much easier. And successes can occasionally be brought into the larger Wikipedias such as we did with Our World in Data interactive graphs...
https://en.wikipedia.org/wiki/Wheat#Production_and_consumption
We have now added 100s of these in various languages. And they are getting thousands of plays a day.
James
On Sat, Jul 11, 2026 at 3:28 PM Charles Roberson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles
On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on
facts (i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot
takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different
choices (i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other
authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the
experience and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps
(i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across
5 models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn
tend to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant
differences in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain
of more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or
"addiction" rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into
"local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l > wikimedia-l@lists.wikimedia.org wrote: > > > Yah a search engine that actually gives real references that > supports the statements in question would be amazing. > > Almost like .. a Knowledge Engine. ;-) >
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through > acquisition, and even if that was financially feasible, you'd begin > by > inheriting a lot of corporate practices that aren't really consistent > with Wikimedia values. > > But perhaps there is a middle ground where Wikimedia seeks to define > more clearly the terms of engagement that it wants with search > engines > (clear attribution, clear and correct references, calls-to-edit, > etc.), and then finds and recognizes search partners who implement > those. To Luis' point, that need not be done by WMF. > > Warmly, > > Erik > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
The blue "Return to article" button works just fine. But yes thanks for pointing out that the back button within the Google chrome browser does not work. Will work on fixing that.
J
On Sat, Jul 11, 2026 at 4:54 PM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Are you talking about that godawful thing that I just clicked on in the "wheat" article, that the back button doesn't work on?
I'm pulling that out. Sorry, but "back button works" is a basic thing of Web functionality. I should not need to click a "return to article" button to get back where I was.
Todd
On Sat, Jul 11, 2026 at 8:32 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
English Wikipedia is a conservative organization / project as are many other large versions of Wikipedia. Many stakeholders need to be convinced / brought on board to make even relatively minor changes. Innovating within smaller versions of Wikipedia or in other projects outside Wikipedia is much easier. And successes can occasionally be brought into the larger Wikipedias such as we did with Our World in Data interactive graphs...
https://en.wikipedia.org/wiki/Wheat#Production_and_consumption
We have now added 100s of these in various languages. And they are getting thousands of plays a day.
James
On Sat, Jul 11, 2026 at 3:28 PM Charles Roberson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles
On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on
facts (i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot
takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different
choices (i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other
authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the
experience and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps
(i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across
5 models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn
tend to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant
differences in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a
chain of more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or
"addiction" rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into
"local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> > > On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < > wikimedia-l@lists.wikimedia.org> wrote: > >> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l >> wikimedia-l@lists.wikimedia.org wrote: >> >> > Yah a search engine that actually gives real references that >> supports the statements in question would be amazing. >> >> Almost like .. a Knowledge Engine. ;-) >> > > Is the WMF building an MCP server to connect Wikipedia and Wikidata > directly to Gemini, Claude, and ChatGPT? This is a more lightweight, > backdoor way to leverage the audience of those platforms but present > structured outputs to AI chats for users based on Wikimedia knowledge. If > we did so, we could present citations within the returned responses to > users, and those platforms make it transparent to the user when they are > calling a particular tool. > > Steven Walling > > Sadly, the only realistic path I see there would be through >> acquisition, and even if that was financially feasible, you'd begin >> by >> inheriting a lot of corporate practices that aren't really >> consistent >> with Wikimedia values. >> >> But perhaps there is a middle ground where Wikimedia seeks to define >> more clearly the terms of engagement that it wants with search >> engines >> (clear attribution, clear and correct references, calls-to-edit, >> etc.), and then finds and recognizes search partners who implement >> those. To Luis' point, that need not be done by WMF. >> >> Warmly, >> >> Erik >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >> guidelines at: >> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to >> wikimedia-l-leave@lists.wikimedia.org > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hey Todd, James and Sage (referring to his recent post about adding video to interfaces)
Your comments highlight exatly where we actually very much are in concensus: Wikipedia is first and foremost curated knowledge by humans for humans (this is the core of my essay here: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed)
I am not arguing that we should have AI writing, or that AI is the primary audience for our content, but rather whether we are entering an era of Zero Click Google, and if the big tech companies see through their vision (i.e. implementaiton of sloppy, adhoc commercially controlled interfaces on top of the open internet i.e.: https://blog.google/products-and-platforms/products/search/search-io-2026/ ) we need to think about this as a distribution channel not an invisible enemy. We cannot afford to miss this modality change in the same way we missed other big changes on the internet. such as social media and the TikTokified enshiftification towards endless scrolling, because we had a guaranteed distribution channel: Google.
Some more thoughts (and realizing that this should be an op-ed/blogpost somewhere).
*Interfaces are cheap, informed curators are expensive*
Agentic AI, Coding tools and LLMs are making the cost of new interfaces *extremely cheap, *so cheap that I am commissioning a complex knowledge repository for a fraction of the cost and time it would take otherwise. Interface projects like WikiProject Med's offline medical Wikipedia App, which used to take several months of highly specialized software development, can now be spun up in a long weekend with a Claude Code Max subscription. Sage's tool is a perfect example of this: perfect for a small market of users, unlikely to be a "headline" Wikimedia tactic for getting in front of users, because Youtube and AI search interfaces already do this exact thing.
What is not cheap for all of the other platforms (and is often paid for by ads), but which we have in abundance, is motivated humans who can continue curating the knowledge (and more than 25 years of experiments on facilitating knowledge equity focused gap filling). The questions we need to address are:
- Do the curators understand the future of distribution to other humans we need to be building for across multiple future internet scenerios? - Do the curation practices serve diverse forms of access (languages, geographies, topics) from that distribution? - Do the curators understand demand and how curation choices affect distribution to that demand? - Can we recruit the next generation of curators who don't "assume" that the pageview metric is the reason we contribute?
*We need to see how distribution to humans is changing *
Our metrics infrastructure overemphasizes our historical abundance of pageviews—80% of which were driven by Google-- without seeing our distribution to these other platforms. We need to *see the distribution* in order to make any decisions about our curatorial practices. .
Once we understand the distribution, we will also see that AI tools building interfaces need more than just access (i.e. Enterprise API or an open license and scraper access), they require knowledge formatting and organization practices (i.e. markdown files, vectorized search interfaces, AI skills, content chunking that makes the RAG search step easier, etc) that necessarily require the content curators to change some of their workflows and practices (emphasizing the authority of original authors in citations, metadata on kind of "human questions" a section describes, etc). Waiting for a handful of people at the Foundation to figure out which of these content organization tactics are important is just not feasible—we, as the curators, need to *as the curators* imagine this future and implement it in our content updates (especially when WMF's payroll relies on the nostalgic pageview-to-fundraising business model).
*We need a Wikimedia specific strategy for the future, not copy our peers*
Most websites are dividing the "content curation" from the representation layer (modern headless CMS's https://en.wikipedia.org/wiki/Headless_content_management_system are increasingly the go to for other publishing websites). Demanding that the Foundation (or Wikimedia projects) maintain our interface as the source of reader interactions is likely not sustainable, or consistent with the way in which knowledge curation now works on the internet.
Our peers in textual content curation implemented radically different strategies 3-5 years ago that build on this seperation of curation and consumption:
- New York Times doubled down on a "captive in an App" strategy because it was a pillar of their approach—a strategy many news organizations adopted as well. This approach is highly inappropriate for the Wikimedia model; we have abundant research showing that our users seek "public service utility" content from us, not timely or trustworthy content. - Britanica has shifted towards an education-market-first model that allows them to build interfaces appropriate to educational needs and garuntee a pipeline of funding. - Reddit optimized for AI tools answering constructive user question -- this also is not our goal, we are curators not "authorities" for answers. - Companies like Healthline optimized for SEO optimization, which gravitates for "generic assumptions of public search" instead of high quality verfiable content (I have found misinformation on healthline multiple times).
My question is: What is the Wikimedia specific business model that allows our curated content to reach the humans we want to reach in this new distribution environment dictated by AI interfaces like RAG and slop-Apps ?
On Sat, Jul 11, 2026 at 12:04 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The blue "Return to article" button works just fine. But yes thanks for pointing out that the back button within the Google chrome browser does not work. Will work on fixing that.
J
On Sat, Jul 11, 2026 at 4:54 PM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Are you talking about that godawful thing that I just clicked on in the "wheat" article, that the back button doesn't work on?
I'm pulling that out. Sorry, but "back button works" is a basic thing of Web functionality. I should not need to click a "return to article" button to get back where I was.
Todd
On Sat, Jul 11, 2026 at 8:32 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
English Wikipedia is a conservative organization / project as are many other large versions of Wikipedia. Many stakeholders need to be convinced / brought on board to make even relatively minor changes. Innovating within smaller versions of Wikipedia or in other projects outside Wikipedia is much easier. And successes can occasionally be brought into the larger Wikipedias such as we did with Our World in Data interactive graphs...
https://en.wikipedia.org/wiki/Wheat#Production_and_consumption
We have now added 100s of these in various languages. And they are getting thousands of plays a day.
James
On Sat, Jul 11, 2026 at 3:28 PM Charles Roberson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles
On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> Forking this conversation, because I don't think we have a shared > framing of what we are competing for in a "Google" Zero landscape dominated > by ChatBot/AI search style RAG citations (i.e. > https://en.wikipedia.org/wiki/Retrieval-augmented_generation). > > For the last 9 months, I've been examining how civil society content > should be showing up in AI search, and there is a missing perspective in > our "we will build it and they will come" approach to Wikipedia. I don't > think we can wait for the big tech companies or European regulatory bodies > to adopt a different idea of how RAG should work. Here is my take from > what I have been engaging with in the AIO/AEO/SEO space: > > *Wikipedia doesn't have content that is SEO/AEO optimized* > > Part of the problem, even if we did have an MCP server is that the > models (at least in my tracking), are pushing many citations away from > "factual" websites, towards "authoritative" websites. This authoritative > content includes: > > - Expert original, analysis that makes strong claims based on > facts (i.e. blog posts by authoritative companies or recently published > ScienceDirect articles) > - Content that has been updated recently, with the biggest "hot > takes" (i.e. I have monitored a couple of prompt pools where citations > shift to newer content after 2-3 months) > - Content that helps users make a decision between different > choices (i.e. review websites, etc) > > > This is following Google's longer-term push towards "human centered > and useful" content (sometimes called E-E-A-T an abbreviation of > experience, expertise, authoritativeness, and trustworthiness, in SEO > world). > https://developers.google.com/search/docs/fundamentals/creating-helpful-cont... > > > To win in an AI optimization battle -- its less about Wikipedia > doing well in the keyword search indexes that led to our content being > visible (which is why we have a reputation as a "fact checking" website) > and more about "winning" in the criteria for what makes a good RAG citation > --- and our content format, is the exact opposite of the EAAT criteria: > > - Wikipedia is not authoriative, but rather points to other > authorities > - We ground our content in anonymity instead of named experts or > instutional process/opinoin > - We rarely do original analysis instead summarizing the > experience and expertise of others, > - Alot of our content is out of date, and self-aware of its gaps > (i.e. maintenance tags), so also is likely to be undermining its own > trustworthiness > > > *All the data points to us being used, but without an official > roundup we are all talking in the dark about different assumed reputation > losses* > > RAG unlike Google Search Indexing, seems to be using Wikipedia for a > fraction of a fraction of responses, favoring these other kinds of > sources: > > - Only 5% of AI overviews have Wikipedia in them: > https://ahrefs.com/blog/most-cited-domains-ai-overviews/ > - I have access to SERanking's corpus of prompt monitoring > across 5 models (ChatGPT, Perplexity, Google Models and they suggest that > in May ~16% of prompts included Wikipdia, and in their most recent month > (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in > commercial AIO data -- so could have gaps. > - Studies from earlier in the year put Wikipedia at about ~13% > of CHATGPT citations ( > https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... > but chatgpt on average includes >20 sources in a response, compared to > googles 5-10 and doesn't expose it in the interface very well) > - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn > tend to represent a greater % of content (in the SERanking data pool nearly > 30% of responses had a Youtube Video cited for instance) > - Domain specific citation pools have pretty significant > differences in "which" sources are being called, with Wikipedia doing well > on some prompt pools: https://generativepulse.ai/report/ > > > > > *RAG/AI search optimization focuses more on intent than keywords, > and we aren't very effective at serving intent, and we don't know where our > optimization options are* > What we need is an understanding of "which actual user reader > behavior are we seeking to serve?". In the past we were extremely lazy, > because keyword search always delivered Wikipedia as "a first". Now we need > our content to be more optimized for the kind of user curiosity driving > their use of a chatbot/search tool: > > - What percentage of prompts or AI searches are informational vs > opinion forming? Are we even a competitor for grounding opinion based > questions or only the informational ones? > - How many of the interactions are two or three steps down a > chain of more "specific" interactions with the chatbot and thus no longer > need "general knowledge" information from Wikipedia, but rather the kinds > of stuff that we rely on our citations to provide ? > - How much are the AI companies optimizing for "sales" or > "addiction" rather than for leading users to reliable content? (I was > tracking a series of informational topics about food that (on ChatGPT and > Google), kept wanting me to continue the conversation by *inviting > me to go to local hamburger restraunts)*. Do we even have a > reasonable chance to be in those searches? > - How much is geolocation forcing more and more responses into > "local" sources rather than "global" websites? In one dataset I tracked, in > Global South countries citations were overwhelmingly to Facebook and > Instagram despite more authoritative academic, news and Wikipedia-type > sites in the same searches from the UK. > > > *We may need to radically change the "readable signals" on our > content pages, meaning changing the Manual of Style, Editing Practices, and > AI enabled enrichment.* > > If we are trying to market Wikipedia's content into AI interfaces, > we also can't do what most AI optimization/marketing agencies would > suggest: writing listicle/FAQ type content that closely matches the > user-queries that folks are giving the IA models (i.e. analysis like: > https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ > ). > > We would then have to experiment with other content types, that _no > longer look like the encyclopedia_. Or we would need to be reconfiguring > the Encyclopedic content to expose enrichments to paragraphs or sections > within the encyclopedia that pretty radically change editorial assupmtions > and our Manual of Style (i.e. instead of simple 1-2 word section headings, > like "History" we may need intent-focused headings like "What is the > history of [x topic]?). > > If we want to compete in the shifting AI search landscape -- we > would need a lot more data from the Foundation on where we are succeeding > or not, and then consider *_radically different_ *ways of exposing > our content in terms of treating RAG systems as a user that needs correct > paths to Wikipedia pages. > > However, this doesn't necessarily need to change the *human reader > experience *, but would need to be about configuring the content > (beyond an MCP server or Enterpise APIs) *for an AI > audience/consumer experience -- *which I haven't seen addressed in > any WMF publications or community conversations. Without a firm theory of > "What kind of consumer is an AI search agent/RAG index?" and "How does our > content need to serve that AI audience?" the editing community won't be > able to adjust its editing practices or weigh in on feature recommendations > that make our content "AI useful". > > As I have written elsewhere, I think there is a inherent audience > for editing/using the Wikis organically: > https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed > -- but its a different question than competing "with other information > sources" for AI as an audience. > > > > > > On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < > wikimedia-l@lists.wikimedia.org> wrote: > >> >> >> On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < >> wikimedia-l@lists.wikimedia.org> wrote: >> >>> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l >>> wikimedia-l@lists.wikimedia.org wrote: >>> >>> > Yah a search engine that actually gives real references that >>> supports the statements in question would be amazing. >>> >>> Almost like .. a Knowledge Engine. ;-) >>> >> >> Is the WMF building an MCP server to connect Wikipedia and Wikidata >> directly to Gemini, Claude, and ChatGPT? This is a more lightweight, >> backdoor way to leverage the audience of those platforms but present >> structured outputs to AI chats for users based on Wikimedia knowledge. If >> we did so, we could present citations within the returned responses to >> users, and those platforms make it transparent to the user when they are >> calling a particular tool. >> >> Steven Walling >> >> Sadly, the only realistic path I see there would be through >>> acquisition, and even if that was financially feasible, you'd >>> begin by >>> inheriting a lot of corporate practices that aren't really >>> consistent >>> with Wikimedia values. >>> >>> But perhaps there is a middle ground where Wikimedia seeks to >>> define >>> more clearly the terms of engagement that it wants with search >>> engines >>> (clear attribution, clear and correct references, calls-to-edit, >>> etc.), and then finds and recognizes search partners who implement >>> those. To Luis' point, that need not be done by WMF. >>> >>> Warmly, >>> >>> Erik >>> _______________________________________________ >>> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >>> guidelines at: >>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>> https://meta.wikimedia.org/wiki/Wikimedia-l >>> Public archives at >>> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >>> To unsubscribe send an email to >>> wikimedia-l-leave@lists.wikimedia.org >> >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >> guidelines at: >> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to >> wikimedia-l-leave@lists.wikimedia.org > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
"My question is: What is the Wikimedia specific business model that allows our curated content to reach the humans we want to reach in this new distribution environment dictated by AI interfaces like RAG and slop-Apps?"
Don't let them dictate. We write Wikipedia like we always have, by people for people. We've never cared about any "business model" and we shouldn't start to today. We just do our thing.
Todd
On Sat, Jul 11, 2026 at 11:44 AM Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hey Todd, James and Sage (referring to his recent post about adding video to interfaces)
Your comments highlight exatly where we actually very much are in concensus: Wikipedia is first and foremost curated knowledge by humans for humans (this is the core of my essay here: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed )
I am not arguing that we should have AI writing, or that AI is the primary audience for our content, but rather whether we are entering an era of Zero Click Google, and if the big tech companies see through their vision (i.e. implementaiton of sloppy, adhoc commercially controlled interfaces on top of the open internet i.e.: https://blog.google/products-and-platforms/products/search/search-io-2026/ ) we need to think about this as a distribution channel not an invisible enemy. We cannot afford to miss this modality change in the same way we missed other big changes on the internet. such as social media and the TikTokified enshiftification towards endless scrolling, because we had a guaranteed distribution channel: Google.
Some more thoughts (and realizing that this should be an op-ed/blogpost somewhere).
*Interfaces are cheap, informed curators are expensive*
Agentic AI, Coding tools and LLMs are making the cost of new interfaces *extremely cheap, *so cheap that I am commissioning a complex knowledge repository for a fraction of the cost and time it would take otherwise. Interface projects like WikiProject Med's offline medical Wikipedia App, which used to take several months of highly specialized software development, can now be spun up in a long weekend with a Claude Code Max subscription. Sage's tool is a perfect example of this: perfect for a small market of users, unlikely to be a "headline" Wikimedia tactic for getting in front of users, because Youtube and AI search interfaces already do this exact thing.
What is not cheap for all of the other platforms (and is often paid for by ads), but which we have in abundance, is motivated humans who can continue curating the knowledge (and more than 25 years of experiments on facilitating knowledge equity focused gap filling). The questions we need to address are:
- Do the curators understand the future of distribution to other
humans we need to be building for across multiple future internet scenerios?
- Do the curation practices serve diverse forms of access (languages,
geographies, topics) from that distribution?
- Do the curators understand demand and how curation choices affect
distribution to that demand?
- Can we recruit the next generation of curators who don't "assume"
that the pageview metric is the reason we contribute?
*We need to see how distribution to humans is changing *
Our metrics infrastructure overemphasizes our historical abundance of pageviews—80% of which were driven by Google-- without seeing our distribution to these other platforms. We need to *see the distribution* in order to make any decisions about our curatorial practices. .
Once we understand the distribution, we will also see that AI tools building interfaces need more than just access (i.e. Enterprise API or an open license and scraper access), they require knowledge formatting and organization practices (i.e. markdown files, vectorized search interfaces, AI skills, content chunking that makes the RAG search step easier, etc) that necessarily require the content curators to change some of their workflows and practices (emphasizing the authority of original authors in citations, metadata on kind of "human questions" a section describes, etc). Waiting for a handful of people at the Foundation to figure out which of these content organization tactics are important is just not feasible—we, as the curators, need to *as the curators* imagine this future and implement it in our content updates (especially when WMF's payroll relies on the nostalgic pageview-to-fundraising business model).
*We need a Wikimedia specific strategy for the future, not copy our peers*
Most websites are dividing the "content curation" from the representation layer (modern headless CMS's https://en.wikipedia.org/wiki/Headless_content_management_system are increasingly the go to for other publishing websites). Demanding that the Foundation (or Wikimedia projects) maintain our interface as the source of reader interactions is likely not sustainable, or consistent with the way in which knowledge curation now works on the internet.
Our peers in textual content curation implemented radically different strategies 3-5 years ago that build on this seperation of curation and consumption:
- New York Times doubled down on a "captive in an App" strategy
because it was a pillar of their approach—a strategy many news organizations adopted as well. This approach is highly inappropriate for the Wikimedia model; we have abundant research showing that our users seek "public service utility" content from us, not timely or trustworthy content.
- Britanica has shifted towards an education-market-first model that
allows them to build interfaces appropriate to educational needs and garuntee a pipeline of funding.
- Reddit optimized for AI tools answering constructive user question
-- this also is not our goal, we are curators not "authorities" for answers.
- Companies like Healthline optimized for SEO optimization, which
gravitates for "generic assumptions of public search" instead of high quality verfiable content (I have found misinformation on healthline multiple times).
My question is: What is the Wikimedia specific business model that allows our curated content to reach the humans we want to reach in this new distribution environment dictated by AI interfaces like RAG and slop-Apps ?
On Sat, Jul 11, 2026 at 12:04 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The blue "Return to article" button works just fine. But yes thanks for pointing out that the back button within the Google chrome browser does not work. Will work on fixing that.
J
On Sat, Jul 11, 2026 at 4:54 PM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Are you talking about that godawful thing that I just clicked on in the "wheat" article, that the back button doesn't work on?
I'm pulling that out. Sorry, but "back button works" is a basic thing of Web functionality. I should not need to click a "return to article" button to get back where I was.
Todd
On Sat, Jul 11, 2026 at 8:32 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
English Wikipedia is a conservative organization / project as are many other large versions of Wikipedia. Many stakeholders need to be convinced / brought on board to make even relatively minor changes. Innovating within smaller versions of Wikipedia or in other projects outside Wikipedia is much easier. And successes can occasionally be brought into the larger Wikipedias such as we did with Our World in Data interactive graphs...
https://en.wikipedia.org/wiki/Wheat#Production_and_consumption
We have now added 100s of these in various languages. And they are getting thousands of plays a day.
James
On Sat, Jul 11, 2026 at 3:28 PM Charles Roberson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Wikimedia-I used to be a premium listserve of high minded ideas about how to move the Foundation into the future. The past few months it has drifted into a lot of whining about AI with no real effort to accomplish anything or even plan to do so.
- Charles
On Sat, Jul 11, 2026 at 1:44 AM Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> What has motivated me to spend time writing Wikipedia over the years > is writing for humans. The fact that the content is openly licensed and the > work is supported by an NGO is also key. > > Personally i do not feel any motivation to write primarily for > trillion dollar machines surrounded by venture capital folks hoping to make > a killing. If the machines want to adapt to human facing content sure. > > The approaches you mention is how Healthline succeeded, they > basically have dozens of articles covering the same topic just addressing > it from a slightly different question. > > J > > > Sent from Gmail Mobile > > On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < > wikimedia-l@lists.wikimedia.org> wrote: > >> Forking this conversation, because I don't think we have a shared >> framing of what we are competing for in a "Google" Zero landscape dominated >> by ChatBot/AI search style RAG citations (i.e. >> https://en.wikipedia.org/wiki/Retrieval-augmented_generation). >> >> For the last 9 months, I've been examining how civil society >> content should be showing up in AI search, and there is a missing >> perspective in our "we will build it and they will come" approach to >> Wikipedia. I don't think we can wait for the big tech companies or European >> regulatory bodies to adopt a different idea of how RAG should work. Here >> is my take from what I have been engaging with in the AIO/AEO/SEO space: >> >> *Wikipedia doesn't have content that is SEO/AEO optimized* >> >> Part of the problem, even if we did have an MCP server is that the >> models (at least in my tracking), are pushing many citations away from >> "factual" websites, towards "authoritative" websites. This authoritative >> content includes: >> >> - Expert original, analysis that makes strong claims based on >> facts (i.e. blog posts by authoritative companies or recently published >> ScienceDirect articles) >> - Content that has been updated recently, with the biggest "hot >> takes" (i.e. I have monitored a couple of prompt pools where citations >> shift to newer content after 2-3 months) >> - Content that helps users make a decision between different >> choices (i.e. review websites, etc) >> >> >> This is following Google's longer-term push towards "human centered >> and useful" content (sometimes called E-E-A-T an abbreviation of >> experience, expertise, authoritativeness, and trustworthiness, in SEO >> world). >> https://developers.google.com/search/docs/fundamentals/creating-helpful-cont... >> >> >> To win in an AI optimization battle -- its less about Wikipedia >> doing well in the keyword search indexes that led to our content being >> visible (which is why we have a reputation as a "fact checking" website) >> and more about "winning" in the criteria for what makes a good RAG citation >> --- and our content format, is the exact opposite of the EAAT criteria: >> >> - Wikipedia is not authoriative, but rather points to other >> authorities >> - We ground our content in anonymity instead of named experts >> or instutional process/opinoin >> - We rarely do original analysis instead summarizing the >> experience and expertise of others, >> - Alot of our content is out of date, and self-aware of its >> gaps (i.e. maintenance tags), so also is likely to be undermining its own >> trustworthiness >> >> >> *All the data points to us being used, but without an official >> roundup we are all talking in the dark about different assumed reputation >> losses* >> >> RAG unlike Google Search Indexing, seems to be using Wikipedia for >> a fraction of a fraction of responses, favoring these other kinds of >> sources: >> >> - Only 5% of AI overviews have Wikipedia in them: >> https://ahrefs.com/blog/most-cited-domains-ai-overviews/ >> - I have access to SERanking's corpus of prompt monitoring >> across 5 models (ChatGPT, Perplexity, Google Models and they suggest that >> in May ~16% of prompts included Wikipdia, and in their most recent month >> (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in >> commercial AIO data -- so could have gaps. >> - Studies from earlier in the year put Wikipedia at about ~13% >> of CHATGPT citations ( >> https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... >> but chatgpt on average includes >20 sources in a response, compared to >> googles 5-10 and doesn't expose it in the interface very well) >> - Comparable "top" Websites, like Youtube, Reddit, and LinkedIn >> tend to represent a greater % of content (in the SERanking data pool nearly >> 30% of responses had a Youtube Video cited for instance) >> - Domain specific citation pools have pretty significant >> differences in "which" sources are being called, with Wikipedia doing well >> on some prompt pools: https://generativepulse.ai/report/ >> >> >> >> >> *RAG/AI search optimization focuses more on intent than keywords, >> and we aren't very effective at serving intent, and we don't know where our >> optimization options are* >> What we need is an understanding of "which actual user reader >> behavior are we seeking to serve?". In the past we were extremely lazy, >> because keyword search always delivered Wikipedia as "a first". Now we need >> our content to be more optimized for the kind of user curiosity driving >> their use of a chatbot/search tool: >> >> - What percentage of prompts or AI searches are informational >> vs opinion forming? Are we even a competitor for grounding opinion based >> questions or only the informational ones? >> - How many of the interactions are two or three steps down a >> chain of more "specific" interactions with the chatbot and thus no longer >> need "general knowledge" information from Wikipedia, but rather the kinds >> of stuff that we rely on our citations to provide ? >> - How much are the AI companies optimizing for "sales" or >> "addiction" rather than for leading users to reliable content? (I was >> tracking a series of informational topics about food that (on ChatGPT and >> Google), kept wanting me to continue the conversation by *inviting >> me to go to local hamburger restraunts)*. Do we even have a >> reasonable chance to be in those searches? >> - How much is geolocation forcing more and more responses into >> "local" sources rather than "global" websites? In one dataset I tracked, in >> Global South countries citations were overwhelmingly to Facebook and >> Instagram despite more authoritative academic, news and Wikipedia-type >> sites in the same searches from the UK. >> >> >> *We may need to radically change the "readable signals" on our >> content pages, meaning changing the Manual of Style, Editing Practices, and >> AI enabled enrichment.* >> >> If we are trying to market Wikipedia's content into AI interfaces, >> we also can't do what most AI optimization/marketing agencies would >> suggest: writing listicle/FAQ type content that closely matches the >> user-queries that folks are giving the IA models (i.e. analysis like: >> https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ >> ). >> >> We would then have to experiment with other content types, that _no >> longer look like the encyclopedia_. Or we would need to be reconfiguring >> the Encyclopedic content to expose enrichments to paragraphs or sections >> within the encyclopedia that pretty radically change editorial assupmtions >> and our Manual of Style (i.e. instead of simple 1-2 word section headings, >> like "History" we may need intent-focused headings like "What is the >> history of [x topic]?). >> >> If we want to compete in the shifting AI search landscape -- we >> would need a lot more data from the Foundation on where we are succeeding >> or not, and then consider *_radically different_ *ways of exposing >> our content in terms of treating RAG systems as a user that needs correct >> paths to Wikipedia pages. >> >> However, this doesn't necessarily need to change the *human reader >> experience *, but would need to be about configuring the content >> (beyond an MCP server or Enterpise APIs) *for an AI >> audience/consumer experience -- *which I haven't seen addressed in >> any WMF publications or community conversations. Without a firm theory of >> "What kind of consumer is an AI search agent/RAG index?" and "How does our >> content need to serve that AI audience?" the editing community won't be >> able to adjust its editing practices or weigh in on feature recommendations >> that make our content "AI useful". >> >> As I have written elsewhere, I think there is a inherent audience >> for editing/using the Wikis organically: >> https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed >> -- but its a different question than competing "with other information >> sources" for AI as an audience. >> >> >> >> >> >> On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < >> wikimedia-l@lists.wikimedia.org> wrote: >> >>> >>> >>> On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < >>> wikimedia-l@lists.wikimedia.org> wrote: >>> >>>> On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l >>>> wikimedia-l@lists.wikimedia.org wrote: >>>> >>>> > Yah a search engine that actually gives real references that >>>> supports the statements in question would be amazing. >>>> >>>> Almost like .. a Knowledge Engine. ;-) >>>> >>> >>> Is the WMF building an MCP server to connect Wikipedia and >>> Wikidata directly to Gemini, Claude, and ChatGPT? This is a more >>> lightweight, backdoor way to leverage the audience of those platforms but >>> present structured outputs to AI chats for users based on Wikimedia >>> knowledge. If we did so, we could present citations within the >>> returned responses to users, and those platforms make it transparent to the >>> user when they are calling a particular tool. >>> >>> Steven Walling >>> >>> Sadly, the only realistic path I see there would be through >>>> acquisition, and even if that was financially feasible, you'd >>>> begin by >>>> inheriting a lot of corporate practices that aren't really >>>> consistent >>>> with Wikimedia values. >>>> >>>> But perhaps there is a middle ground where Wikimedia seeks to >>>> define >>>> more clearly the terms of engagement that it wants with search >>>> engines >>>> (clear attribution, clear and correct references, calls-to-edit, >>>> etc.), and then finds and recognizes search partners who implement >>>> those. To Luis' point, that need not be done by WMF. >>>> >>>> Warmly, >>>> >>>> Erik >>>> _______________________________________________ >>>> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >>>> guidelines at: >>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>> Public archives at >>>> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >>>> To unsubscribe send an email to >>>> wikimedia-l-leave@lists.wikimedia.org >>> >>> _______________________________________________ >>> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >>> guidelines at: >>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>> https://meta.wikimedia.org/wiki/Wikimedia-l >>> Public archives at >>> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >>> To unsubscribe send an email to >>> wikimedia-l-leave@lists.wikimedia.org >> >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >> guidelines at: >> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to >> wikimedia-l-leave@lists.wikimedia.org > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
On 7/11/2026 10:43 AM, Alex Stinson via Wikimedia-l wrote:
*Interfaces are cheap, informed curators are expensive*
Agentic AI, Coding tools and LLMs are making the cost of new interfaces /extremely cheap, /so cheap that I am commissioning a complex knowledge repository for a fraction of the cost and time it would take otherwise. Interface projects like WikiProject Med's offline medical Wikipedia App, which used to take several months of highly specialized software development, can now be spun up in a long weekend with a Claude Code Max subscription. Sage's tool is a perfect example of this: perfect for a small market of users, unlikely to be a "headline" Wikimedia tactic for getting in front of users, because Youtube and AI search interfaces already do this exact thing.
What is not cheap for all of the other platforms (and is often paid for by ads), but which we have in abundance, is motivated humans who can continue curating the knowledge (and more than 25 years of experiments on facilitating knowledge equity focused gap filling). The questions we need to address are:
- Do the curators understand the future of distribution to other humans we need to be building for across multiple future internet scenerios?
- Do the curation practices serve diverse forms of access (languages, geographies, topics) from that distribution?
- Do the curators understand demand and how curation choices affect distribution to that demand?
These three questions all implicitly require some kind of training and education for curators, on matters not directly affecting curation activity. Our volunteers are motivated, yes, but what motivates is the curation they enjoy, not necessarily tangential considerations. People contribute because "I see a typo", "I'm collecting bits of trivia in this space", "This leaves out a key fact", or "That's just flat wrong", and the wiki makes them all (relatively) easy to fix. For us, it's "cheap" as in we've lowered the investment required to participate. Some curators will be willing to invest additional engagement on big-picture issues, so it may nevertheless be helpful to offer such training, but there's still investment cost on both sides and I do not see it moving the needle when focusing on the competitive landscape.
- Can we recruit the next generation of curators who don't "assume" that the pageview metric is the reason we contribute?
I question whether that assumption is in fact prevalent. Even in the world of professional media, those paid to curate or generate content often resist the implications, whether out of principled beliefs or because they fear impact to their livelihood. This conversation has emphasized such metrics because those participating care about reaching a broader audience. But in terms of what motivates curation at a basic level, once it is established that there is some audience out there, I don't know that the order of its magnitude matters all that much. Audience demand results in more information to curate and more space in which to work, and I think that motivates equally with reach if not more. People may find it as satisfying to work on en:Wedding of Taylor Swift and Travis Kelce, even though en:Taylor Swift and en:Travis Kelce undoubtedly have higher pageviews.
To use the wheat production infographic as another illustration, personally I care less about the interface behavior than what I could do to help better curate. Suppose I had information about how much wheat Sudan produced in 2011 (currently "no data"), or wanted to address the anachronistic use of present-day national boundaries going back all the way to 1961? Where would I even start? Otherwise, maybe it increases the visual appeal of one article a bit, and we could even reproduce that across a bunch of articles, but we'd also be throwing up massive barriers to maintenance.
What we most need is to simplify, to make it easier for volunteers to steer in the direction we would like them to go. Then they can self-select according to what meets their interests. I believe the success of Wikidata so far - "success" in the sense that its potential aligns more with "early Wikipedia" than "early Wikiversity" - is in large part because it provides a vast new space for curation, and has been integrated in a way that makes it easy for those interested to go and participate, while it is simultaneously easy to ignore for those who are less interested.
--Michael Snow
Hoi, A few questions:
- how expensive is it to add all references, undoubled, from all our projects to Wikidata AND have them link to where they are used - how expensive is it to check if these references are still there and link to content at archive.org - how expensive is it to create a search engine with the subjects and the titles what can be searched for - consider this an experiment, what would happen if we make it available to the public without fanfare? - how expensive would it be to allow a public to add links to youtube and whatever when they are linked to subjects [1]
- how expensive is it to flag to Wikipedias and other projects that lists are likely incorrect based on information that we have in the wiki sphere?
[1] I watched How gut microbes keep us healthy as we age | Tim Spector & Nicola Segata - YouTube https://www.youtube.com/watch?v=d_3dfXbVl2k could link it to these two scientists but I am adding citations for the key paper.. Gut micro-organisms associated with health, nutrition and dietary interventions - Scholia https://qlever.scholia.wiki/work/Q140511862
Thanks, GerardM
On Sat, 11 Jul 2026 at 22:59, Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On 7/11/2026 10:43 AM, Alex Stinson via Wikimedia-l wrote:
*Interfaces are cheap, informed curators are expensive*
Agentic AI, Coding tools and LLMs are making the cost of new interfaces *extremely cheap, *so cheap that I am commissioning a complex knowledge repository for a fraction of the cost and time it would take otherwise. Interface projects like WikiProject Med's offline medical Wikipedia App, which used to take several months of highly specialized software development, can now be spun up in a long weekend with a Claude Code Max subscription. Sage's tool is a perfect example of this: perfect for a small market of users, unlikely to be a "headline" Wikimedia tactic for getting in front of users, because Youtube and AI search interfaces already do this exact thing.
What is not cheap for all of the other platforms (and is often paid for by ads), but which we have in abundance, is motivated humans who can continue curating the knowledge (and more than 25 years of experiments on facilitating knowledge equity focused gap filling). The questions we need to address are:
- Do the curators understand the future of distribution to other
humans we need to be building for across multiple future internet scenerios?
- Do the curation practices serve diverse forms of access (languages,
geographies, topics) from that distribution?
- Do the curators understand demand and how curation choices affect
distribution to that demand?
These three questions all implicitly require some kind of training and education for curators, on matters not directly affecting curation activity. Our volunteers are motivated, yes, but what motivates is the curation they enjoy, not necessarily tangential considerations. People contribute because "I see a typo", "I'm collecting bits of trivia in this space", "This leaves out a key fact", or "That's just flat wrong", and the wiki makes them all (relatively) easy to fix. For us, it's "cheap" as in we've lowered the investment required to participate. Some curators will be willing to invest additional engagement on big-picture issues, so it may nevertheless be helpful to offer such training, but there's still investment cost on both sides and I do not see it moving the needle when focusing on the competitive landscape.
- Can we recruit the next generation of curators who don't "assume"
that the pageview metric is the reason we contribute?
I question whether that assumption is in fact prevalent. Even in the world of professional media, those paid to curate or generate content often resist the implications, whether out of principled beliefs or because they fear impact to their livelihood. This conversation has emphasized such metrics because those participating care about reaching a broader audience. But in terms of what motivates curation at a basic level, once it is established that there is some audience out there, I don't know that the order of its magnitude matters all that much. Audience demand results in more information to curate and more space in which to work, and I think that motivates equally with reach if not more. People may find it as satisfying to work on en:Wedding of Taylor Swift and Travis Kelce, even though en:Taylor Swift and en:Travis Kelce undoubtedly have higher pageviews.
To use the wheat production infographic as another illustration, personally I care less about the interface behavior than what I could do to help better curate. Suppose I had information about how much wheat Sudan produced in 2011 (currently "no data"), or wanted to address the anachronistic use of present-day national boundaries going back all the way to 1961? Where would I even start? Otherwise, maybe it increases the visual appeal of one article a bit, and we could even reproduce that across a bunch of articles, but we'd also be throwing up massive barriers to maintenance.
What we most need is to simplify, to make it easier for volunteers to steer in the direction we would like them to go. Then they can self-select according to what meets their interests. I believe the success of Wikidata so far - "success" in the sense that its potential aligns more with "early Wikipedia" than "early Wikiversity" - is in large part because it provides a vast new space for curation, and has been integrated in a way that makes it easy for those interested to go and participate, while it is simultaneously easy to ignore for those who are less interested.
--Michael Snow _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
On Sat, Jul 11, 2026 at 1:59 PM Michael Snow via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What we most need is to simplify, to make it easier for volunteers to steer in the direction we would like them to go. Then they can self-select according to what meets their interests. I believe the success of Wikidata so far - "success" in the sense that its potential aligns more with "early Wikipedia" than "early Wikiversity" - is in large part because it provides a vast new space for curation, and has been integrated in a way that makes it easy for those interested to go and participate, while it is simultaneously easy to ignore for those who are less interested.
I think this gets at the heart of it: early Wikipedia was a vast space for curation — with an implicit curation scope of 'anything you've ever read'.
Since then, there's been a generational shift in how people learn about the topics they care about, and how people share what they know for a public audience. Now there's a wide-open curation space for video.
On Sat, Jul 11, 2026 at 10:44 AM Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
*Interfaces are cheap, informed curators are expensive*
[...]
Sage's tool is a perfect example of this: perfect for a small market of users, unlikely to be a "headline" Wikimedia tactic for getting in front of users, because Youtube and AI search interfaces already do this exact thing.
I think the market for a video curation space — especially one that is part of a continuum with our encyclopedia, putting reliable knowledge sources in context and in conversation with each other, moving ever closer to solid core of shared knowledge — is not small. YouTube (and other video platforms) are missing exactly this kind of curation space. They are built around serving you new content, and keeping the creators on the content grind, not curating the best of what's already there. If I want to learn about a topic, I don't want every video that matches my search. I want the best videos — selected by someone who has watched a lot more videos about that topic than I have, from the most reliable knowers about the topic. If there's something inconsistent or incorrect, or if it represents a particular perspective that isn't widely accepted, I want to see context from someone who curated it in the same way a Wikipedia editor curates the text sources that go into a Wikipedia article.
I think the audience for video curation is — to oversimplify a bit — my kids' whole generation. My kids don't made videos, but for the topics they're interested in, they would do a lot better at curating videos than I would... and it would cover a whole lot of knowledge that isn't part of the traditional text ecosystem.
So far https://wikiplus.video/ is a Wiki Education side project; we're excited about it, but not quite sure where we want to go with it. But it's already scratching an itch for me personally — I start there instead of Wikipedia when I want to read about something. I'm still first-and-foremost text-oriented, but I keep being surprised at just how much great shared knowledge there is in niche YouTube videos about (some aspect of) an encyclopedic topic.
-Sage Wiki Education
What do you mean by "we have succeeded"?
-- Christophe
On Sat, 11 Jul 2026 at 07:44, Todd Allen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
We sure do know what a good AI strategy looks like for Wikipedia. No AI on Wikipedia.
We have not succeeded by being FaceGramTwitTube. We have succeeded by not being like them.
So, same here. No "latest and greatest". No AI on Wikipedia. Ever, for any reason, period. Wikipedia is written by people for people.
Todd
On Fri, Jul 10, 2026 at 11:09 PM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
What has motivated me to spend time writing Wikipedia over the years is writing for humans. The fact that the content is openly licensed and the work is supported by an NGO is also key.
Personally i do not feel any motivation to write primarily for trillion dollar machines surrounded by venture capital folks hoping to make a killing. If the machines want to adapt to human facing content sure.
The approaches you mention is how Healthline succeeded, they basically have dozens of articles covering the same topic just addressing it from a slightly different question.
J
Sent from Gmail Mobile
On Fri, Jul 10, 2026 at 21:24 Alex Stinson via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Forking this conversation, because I don't think we have a shared framing of what we are competing for in a "Google" Zero landscape dominated by ChatBot/AI search style RAG citations (i.e. https://en.wikipedia.org/wiki/Retrieval-augmented_generation).
For the last 9 months, I've been examining how civil society content should be showing up in AI search, and there is a missing perspective in our "we will build it and they will come" approach to Wikipedia. I don't think we can wait for the big tech companies or European regulatory bodies to adopt a different idea of how RAG should work. Here is my take from what I have been engaging with in the AIO/AEO/SEO space:
*Wikipedia doesn't have content that is SEO/AEO optimized*
Part of the problem, even if we did have an MCP server is that the models (at least in my tracking), are pushing many citations away from "factual" websites, towards "authoritative" websites. This authoritative content includes:
- Expert original, analysis that makes strong claims based on facts
(i.e. blog posts by authoritative companies or recently published ScienceDirect articles)
- Content that has been updated recently, with the biggest "hot
takes" (i.e. I have monitored a couple of prompt pools where citations shift to newer content after 2-3 months)
- Content that helps users make a decision between different choices
(i.e. review websites, etc)
This is following Google's longer-term push towards "human centered and useful" content (sometimes called E-E-A-T an abbreviation of experience, expertise, authoritativeness, and trustworthiness, in SEO world). https://developers.google.com/search/docs/fundamentals/creating-helpful-cont...
To win in an AI optimization battle -- its less about Wikipedia doing well in the keyword search indexes that led to our content being visible (which is why we have a reputation as a "fact checking" website) and more about "winning" in the criteria for what makes a good RAG citation --- and our content format, is the exact opposite of the EAAT criteria:
- Wikipedia is not authoriative, but rather points to other
authorities
- We ground our content in anonymity instead of named experts or
instutional process/opinoin
- We rarely do original analysis instead summarizing the experience
and expertise of others,
- Alot of our content is out of date, and self-aware of its gaps
(i.e. maintenance tags), so also is likely to be undermining its own trustworthiness
*All the data points to us being used, but without an official roundup we are all talking in the dark about different assumed reputation losses*
RAG unlike Google Search Indexing, seems to be using Wikipedia for a fraction of a fraction of responses, favoring these other kinds of sources:
- Only 5% of AI overviews have Wikipedia in them:
https://ahrefs.com/blog/most-cited-domains-ai-overviews/
- I have access to SERanking's corpus of prompt monitoring across 5
models (ChatGPT, Perplexity, Google Models and they suggest that in May ~16% of prompts included Wikipdia, and in their most recent month (June), ~13% of prompts. SERankings corpus is probably the # 4 or 5 in commercial AIO data -- so could have gaps.
- Studies from earlier in the year put Wikipedia at about ~13% of
CHATGPT citations ( https://www.prnewswire.com/news-releases/wikipedia-and-reddit-now-drive-over... but chatgpt on average includes >20 sources in a response, compared to googles 5-10 and doesn't expose it in the interface very well)
- Comparable "top" Websites, like Youtube, Reddit, and LinkedIn tend
to represent a greater % of content (in the SERanking data pool nearly 30% of responses had a Youtube Video cited for instance)
- Domain specific citation pools have pretty significant differences
in "which" sources are being called, with Wikipedia doing well on some prompt pools: https://generativepulse.ai/report/
*RAG/AI search optimization focuses more on intent than keywords, and we aren't very effective at serving intent, and we don't know where our optimization options are* What we need is an understanding of "which actual user reader behavior are we seeking to serve?". In the past we were extremely lazy, because keyword search always delivered Wikipedia as "a first". Now we need our content to be more optimized for the kind of user curiosity driving their use of a chatbot/search tool:
- What percentage of prompts or AI searches are informational vs
opinion forming? Are we even a competitor for grounding opinion based questions or only the informational ones?
- How many of the interactions are two or three steps down a chain
of more "specific" interactions with the chatbot and thus no longer need "general knowledge" information from Wikipedia, but rather the kinds of stuff that we rely on our citations to provide ?
- How much are the AI companies optimizing for "sales" or
"addiction" rather than for leading users to reliable content? (I was tracking a series of informational topics about food that (on ChatGPT and Google), kept wanting me to continue the conversation by *inviting me to go to local hamburger restraunts)*. Do we even have a reasonable chance to be in those searches?
- How much is geolocation forcing more and more responses into
"local" sources rather than "global" websites? In one dataset I tracked, in Global South countries citations were overwhelmingly to Facebook and Instagram despite more authoritative academic, news and Wikipedia-type sites in the same searches from the UK.
*We may need to radically change the "readable signals" on our content pages, meaning changing the Manual of Style, Editing Practices, and AI enabled enrichment.*
If we are trying to market Wikipedia's content into AI interfaces, we also can't do what most AI optimization/marketing agencies would suggest: writing listicle/FAQ type content that closely matches the user-queries that folks are giving the IA models (i.e. analysis like: https://neilpatel.com/marketing-stats/trust-signals-ai-engines-reward-most/ ).
We would then have to experiment with other content types, that _no longer look like the encyclopedia_. Or we would need to be reconfiguring the Encyclopedic content to expose enrichments to paragraphs or sections within the encyclopedia that pretty radically change editorial assupmtions and our Manual of Style (i.e. instead of simple 1-2 word section headings, like "History" we may need intent-focused headings like "What is the history of [x topic]?).
If we want to compete in the shifting AI search landscape -- we would need a lot more data from the Foundation on where we are succeeding or not, and then consider *_radically different_ *ways of exposing our content in terms of treating RAG systems as a user that needs correct paths to Wikipedia pages.
However, this doesn't necessarily need to change the *human reader experience *, but would need to be about configuring the content (beyond an MCP server or Enterpise APIs) *for an AI audience/consumer experience -- *which I haven't seen addressed in any WMF publications or community conversations. Without a firm theory of "What kind of consumer is an AI search agent/RAG index?" and "How does our content need to serve that AI audience?" the editing community won't be able to adjust its editing practices or weigh in on feature recommendations that make our content "AI useful".
As I have written elsewhere, I think there is a inherent audience for editing/using the Wikis organically: https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-06-21/Op-ed -- but its a different question than competing "with other information sources" for AI as an audience.
On Fri, Jul 10, 2026 at 2:30 PM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 9:52 AM Erik Moeller via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
On Fri, Jul 10, 2026 at 5:59 PM James Heilman via Wikimedia-l wikimedia-l@lists.wikimedia.org wrote:
Yah a search engine that actually gives real references that
supports the statements in question would be amazing.
Almost like .. a Knowledge Engine. ;-)
Is the WMF building an MCP server to connect Wikipedia and Wikidata directly to Gemini, Claude, and ChatGPT? This is a more lightweight, backdoor way to leverage the audience of those platforms but present structured outputs to AI chats for users based on Wikimedia knowledge. If we did so, we could present citations within the returned responses to users, and those platforms make it transparent to the user when they are calling a particular tool.
Steven Walling
Sadly, the only realistic path I see there would be through
acquisition, and even if that was financially feasible, you'd begin by inheriting a lot of corporate practices that aren't really consistent with Wikimedia values.
But perhaps there is a middle ground where Wikimedia seeks to define more clearly the terms of engagement that it wants with search engines (clear attribution, clear and correct references, calls-to-edit, etc.), and then finds and recognizes search partners who implement those. To Luis' point, that need not be done by WMF.
Warmly,
Erik _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Hoi, A charity, the World Cancer Research Fund, is all about the "relation of Diet, Nutrition, Physical Activity and Cancer". When you look up the research they sponsored, you find many epidemiologists, when you then look up their Scholia you typically find an area of research that could use a lot of tender loving care.
They claim three areas of activity; research, influencing cancer professionals and influencing cancer patients. Plenty of papers are already known to Wikidata that prove the point that the occurrence and the outcome of cancer are dramatically influenced.
So when I have the time, I work on scientists like Dr. Dieuwertje Kok. We used to have tools that would import all her works from ORCiD. We could revive them again. When we have a WikiFind, we could check out the content of a website like https://www.wkof.nl/ add their website and link to the Scholias of scientists, like claims to papers and by inference tell people to consider what they can in relation to the 40 often occuring cancers that can be prevented.
This is just a suggestion that we could do this. People work on this kind of data. We should serve it as part of the knowledge that we have. Thanks, GerardM
On Fri, 10 Jul 2026 at 17:59, James Heilman jmh649@gmail.com wrote:
Yah a search engine that actually gives real references that supports the statements in question would be amazing. Google currently gives mostly fake references (ie the references they provide often do not support the content it is attached behind).
J
On Fri, Jul 10, 2026 at 7:20 AM Gerard Meijssen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hoi. Given the lack of maintenance of many Wikipedia articles, eg the median time for a correction of a retracted reference is 3.68 years, the notion that our references will save us is unlikely. Given how big we are on references, we could expose them in our own search engine where we combine what we have, linked to subjects and enable people to add sources and work on our content..
This is not a Wikipedia centred idea and I expect that it will not be considered. But hey, what other ideas are out there that have a chance to expose people to our content? Thanks, GerardM
On Fri, 10 Jul 2026 at 03:26, Kimmo Virtanen via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
Hi,
One note that hasn't been mentioned is that the future UI for consuming information will be "agentic," where the tool used will search for and collect the latest information on the topics. This means systems will prioritize content where the information source is described (i.e., references), as accuracy can be validated from there. So references are one of the strengths of Wikimedia services compared to almost all other sites on the internet. This doesn't really solve the eyeball problem since the content is still used externally, but it should be good for something.
Br, -- Kimmo Virtanen, Zache
On Fri, Jul 10, 2026 at 3:14 AM Steven Walling via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
EU data provenance regulations do not require foundation model providers to link to the constituent sources of AI model knowledge. In fact the disclosure rules don’t even apply to foundation models themselves yet, only services. And they only require disclosing that text or other media were generated by AI (via the interface and metadata). It’s not going to help fix brand awareness and loss in referral traffic, only transparency that a particular piece of content is or is not AI generated.
Google already links to Wikipedia when it is used as a source in search answers, and we’re still seeing a traffic decline. We shouldn’t wait to assume that EU regulations or walled gardens are going to save us from being disintermediated.
The drop in enwiki traffic in particular is likely indicative of a structural shift considering that in just two years (2024-2026) we have gone from 33% of Americans using chatbots and AI summaries to 49% ( https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbo... )
Steven Walling
On Thu, Jul 9, 2026 at 7:39 AM Ilario Valdelli via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
James, thank you for grounding this with real numbers rather than impressions — the Similarweb figure and the cross-language health content analysis are exactly the kind of evidence this thread needs, and they're worth taking as the starting point.
That said, I don't think there's much left to gain from debating what should have been done differently. What happened, happened — the past doesn't come back, and no amount of hindsight changes the trajectory from here.
What matters more now is the immediate ahead of us, and looking at it with a different mindset: not only as a problem, but as something that might open an opportunity — because very often what looks like a threat to us is somebody else's opportunity, and figuring out which side of that split we're on matters more than assigning blame for the past.
A few things, connected, that point in that direction (even not a complete picture). AI infrastructure spending is running above $1 trillion globally this year, with investors increasingly nervous about the return on it. [1] At the same time, open-weight models that run locally — largely out of China, but not only — are now within a few points of frontier benchmarks at a fraction of the cost.[2] Regulation is starting to require actual, checkable data provenance rather than declared compliance (EU AI Act for instance). And large parts of the web are moving toward walling off AI crawlers and charging for metered access by default.[3]
Put together, I think James's numbers might not just describe a decline — they might be the leading edge of a market starting to separate "content that's cheap to grab" from "content whose origin can be verified," and pricing them very differently. If that's right, it matters most exactly where James's data points: health content is where the cost of being wrong is highest, and where verifiable sourcing should be worth the most, not the least.
kind regards
Ilario [1] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-a... [2] https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://computingforgeeks.com/open-source-llm-comparison/ [3 https://about.bnef.com/insights/data-centers/ai-data-center-build-advances-at-full-speed-five-things-to-know/ https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies...
On Thu, Jul 9, 2026 at 10:39 AM James Heilman via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
The question regarding what percent of our traffic comes from search? Similarweb says about 80%, so basically we could lose another 80% of our current traffic.
https://www.similarweb.com/website/wikipedia.org/#traffic-sources
Is the fall in traffic happening in just big languages? No it is happening in many languages of Wikipedia, at least for health content, There were some exceptions like Persian and Chinese that has seen a small degree of growth
https://github.com/nethahussain/Pageviews-in-Medicine/blob/main/analysis/cov...
J
On Thu, Jul 9, 2026 at 10:18 AM Yaroslav Blanter via Wikimedia-l < wikimedia-l@lists.wikimedia.org> wrote:
> I do not think Google Zero will be (is) experienced by anyone except > for us as a problem. > > Sure, the current AI search is not perfect, but it will be much > better on a horizon of months. It is unfortunate that the end users take > the results uncritically rather than checking them clicking on the source > links, but it is what it is. May be this culture would change, but this > timescale is way way longer than the time to improve the AI search quality. > > Except for obvious financial implications (which are very important > though) it looks like Wikipedia will be mainly the base for training AI. It > is ok, in real life most people do not read encyclopaedias. It is important > to be the training base as well, but we should also realize that if this is > our role then some priorities must be shifted. For example, AI do not care > whether the article looks nice or not, they do not care about usability. > They probably do not care about categories and many other things we spent > years trying to make them perfect. They do care about content though. > > On the other hand, there are clearly things AI can not do, which > were already mentioned here. They can not read offline material, and much > of the paywall material. They can not produce pictures out of nothing - > which means Commons might have a very different role from what it has now. > It might even become the flagship project (unlikely though with the current > Commons community). But all these things require attention, and many of > them will be resisted by the community in the first place. > > Best > Yaroslav > > On Thu, Jul 9, 2026 at 9:33 AM Anders Wennersten via Wikimedia-l < > wikimedia-l@lists.wikimedia.org> wrote: > >> I agree we ought to analyze better our situation, before we run to >> conclusions and discuss remedies >> >> Swwp has seen a decrease in traffic like others but in June we have >> seen an increase even if we include only accesses from mobiles (at least >> for June) >> >> >> https://stats.wikimedia.org/#/sv.wikipedia.org/reading/total-page-views/norm... >> >> What we have seen more clearly though is that accesses of the most >> read articles per day in decreasing. >> >> We have also more subjectively from media, and social media seen a >> growing skepticism toward AI-generated results and an increase in the trust >> of Wikipedia results >> >> Our general conclusions this far are >> >> *We are doing well in our core encyclopedic articles (like of >> administrative divisions within a country) , but less well in newslike >> articles (which the internet readers asks more frequently, but where >> newswebsites always been stronger). And we should concentrate our efforts >> where we are strong >> >> *We are in a late stage in our general product life cycle and must >> accept this. We will never have the increase in editors etc like we hade 15 >> yeas ago, but must do what we can do thrive in a this mode. For us - less >> angry internal argument, partially done by avoiding articles of infected >> subjects like Israel-Palestine >> >> And for Google Zero. I believe it is them who has a problem not us. >> Early on, and still, all new PC user were forced into Microsoft Edge and >> Bing which lead us to Firefox and Google. Google zero as I see will lead >> to an exodus från Google, there are many other search engines that works >> fine >> >> Wikipedia is great >> >> Anders >> >> >> >> >> Den 2026-07-09 kl. 08:35, skrev Kimmo Virtanen via Wikimedia-l: >> >> Hi, >> >> Secondly, I'd love to see initiative outside of enwiki. If you look >>> at all the ~1000 Wikimedia projects, almost all of them are seeing an >>> increase in traffic ; only a few big old wikipedias are seeing a decrease >>> (which are also the ones with the most traffic, hence why it's more visible >>> but it obfuscate at the big picture) >> >> >> Based on stats total pageviews of human users of all wikis are >> dropping >> - >> https://stats.wikimedia.org/#/all-projects/reading/total-page-views/normal%7... >> >> >> Br, >> -- Kimmo Virtanen, Zache >> >> On Thu, Jul 9, 2026 at 9:13 AM Nicolas VIGNERON via Wikimedia-l < >> wikimedia-l@lists.wikimedia.org> wrote: >> >>> Hi y'all, >>> >>> First could someone remind me how much traffic comes from Google >>> (I have a vague memory from last year that this is already very low - maybe >>> 10 or 20 % - and that LLMs are already bringing more traffic to Wikipedia, >>> if true we essentially are at Google Zero already). >>> >>> Secondly, I'd love to see initiative outside of enwiki. If you >>> look at all the ~1000 Wikimedia projects, almost all of them are seeing an >>> increase in traffic ; only a few big old wikipedias are seeing a decrease >>> (which are also the ones with the most traffic, hence why it's more visible >>> but it obfuscate at the big picture). >>> I think we should start by looking at and learning from the >>> "successful" projects instead of focusing on the declining ones. For >>> instance, Wikisources always have been badly to not-at-all referenced by >>> Google and yet there is a strong increase in pageviews recently. >>> >>> Cheers, >>> Nicolas >>> >>> Le jeu. 9 juil. 2026 à 03:38, Michael Snow via Wikimedia-l < >>> wikimedia-l@lists.wikimedia.org> a écrit : >>> >>>> On 7/8/2026 12:52 PM, Luis Villa via Wikimedia-l wrote: >>>> > - Scale of impact is important, but it isn’t everything. >>>> Sometimes >>>> > we’ll have to build things without knowing if they’ll scale, or >>>> even >>>> > perhaps knowing that they won’t scale (eg Depths of Wikipedia >>>> is one >>>> > of the best community-builders we have even though it has >>>> “only” a few >>>> > million followers). “That’s good but it won’t scale to enwiki” >>>> can >>>> > never be allowed to be a blocker for anyone. If the experiment >>>> is >>>> > actually great, then we’ll figure out a way to get it onto >>>> enwiki. Or >>>> > not! That shouldn’t stop people. >>>> >>>> At this point, no single initiative (whether community, chapter, >>>> or >>>> WMF-led) will have an inherent capacity to scale to enwiki. To do >>>> so >>>> would require corporate-style resources with a corresponding >>>> top-down >>>> mandate for implementation, and risk (rightly) triggering >>>> resistance to >>>> the point of launching a true durable fork, for once. Instead, >>>> the best >>>> approach would be to find ideas that are good enough to >>>> propagate, >>>> encourage them, and when they have propagated enough, they will >>>> scale >>>> with relatively little intervention needed. >>>> >>>> --Michael Snow >>>> >>>> _______________________________________________ >>>> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >>>> guidelines at: >>>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>>> https://meta.wikimedia.org/wiki/Wikimedia-l >>>> Public archives at >>>> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >>>> To unsubscribe send an email to >>>> wikimedia-l-leave@lists.wikimedia.org >>> >>> _______________________________________________ >>> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >>> guidelines at: >>> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >>> https://meta.wikimedia.org/wiki/Wikimedia-l >>> Public archives at >>> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >>> To unsubscribe send an email to >>> wikimedia-l-leave@lists.wikimedia.org >> >> >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org >> >> _______________________________________________ >> Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, >> guidelines at: >> https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and >> https://meta.wikimedia.org/wiki/Wikimedia-l >> Public archives at >> https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... >> To unsubscribe send an email to >> wikimedia-l-leave@lists.wikimedia.org > > _______________________________________________ > Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, > guidelines at: > https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and > https://meta.wikimedia.org/wiki/Wikimedia-l > Public archives at > https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... > To unsubscribe send an email to > wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- Ilario Valdelli Skype: valdelli Tel: +41764821371 http://www.wikimedia.ch _______________________________________________ Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
Wikimedia-l mailing list -- wikimedia-l@lists.wikimedia.org, guidelines at: https://meta.wikimedia.org/wiki/Mailing_lists/Guidelines and https://meta.wikimedia.org/wiki/Wikimedia-l Public archives at https://lists.wikimedia.org/hyperkitty/list/wikimedia-l@lists.wikimedia.org/... To unsubscribe send an email to wikimedia-l-leave@lists.wikimedia.org
-- James Heilman MD, CCFP-EM, Wikipedian
wikimedia-l@lists.wikimedia.org