Dear all,
We are preparing an academic research project involving the large-scale analysis of art-historical material from Wikimedia Commons, conducted jointly by the University of Marburg and the Getty Research Institute.
For this project, we anticipate requiring access to at least five million Commons images. High-resolution originals are not necessary; thumbnails would suffice.
Before we initiate any large-scale retrieval, we would be grateful if you could advise us on the preferred way to access this material. In particular, we would be grateful for advice on whether an existing bulk dataset provides Commons thumbnails on this scale, or whether a project of this size should make use of, or request, any special API arrangements. As we can identify and filter the relevant image URLs from the metadata in advance, the main question concerns the recommended method for retrieving the images themselves.
If you have any more questions about the project, please do not hesitate to get in touch.
Thank you, and all the best,
Stefanie __________________________
Dr. Stefanie Schneider Wissenschaftliche Mitarbeiterin
Philipps-Universität Marburg Kunstgeschichtliches Institut Biegenstraße 11 35037 Marburg
Raum: 00003 Telefon: +49 6421 28-22174 E-Mail: stefanie.schneider@uni-marburg.de
https://www.uni-marburg.de/de/fb09/khi/dr-stefanie-schneider
Hi Stephanie, Thank you for reaching out!
The short answer is that there is no bulk data set for thumbnails, and the right path depends a little on your timeline for retrieving the images. As a first step, I recommend reading through https://wikitech.wikimedia.org/wiki/Robot_policy to learn more about the existing options and best practices. Should you then still need support or have questions, please feel welcome to reach out to us via the contact email listed in the last section of the policy.
Thanks again, Birgit
On Mon, Aug 10, 2026 at 4:53 PM Stefanie Schneider via Wikitech-l < wikitech-l@lists.wikimedia.org> wrote:
Dear all,
We are preparing an academic research project involving the large-scale analysis of art-historical material from Wikimedia Commons, conducted jointly by the University of Marburg and the Getty Research Institute.
For this project, we anticipate requiring access to at least five million Commons images. High-resolution originals are not necessary; thumbnails would suffice.
Before we initiate any large-scale retrieval, we would be grateful if you could advise us on the preferred way to access this material. In particular, we would be grateful for advice on whether an existing bulk dataset provides Commons thumbnails on this scale, or whether a project of this size should make use of, or request, any special API arrangements. As we can identify and filter the relevant image URLs from the metadata in advance, the main question concerns the recommended method for retrieving the images themselves.
If you have any more questions about the project, please do not hesitate to get in touch.
Thank you, and all the best,
Stefanie __________________________
Dr. Stefanie Schneider Wissenschaftliche Mitarbeiterin
Philipps-Universität Marburg Kunstgeschichtliches Institut Biegenstraße 11 35037 Marburg
Raum: 00003 Telefon: +49 6421 28-22174 E-Mail: stefanie.schneider@uni-marburg.de
https://www.uni-marburg.de/de/fb09/khi/dr-stefanie-schneider _______________________________________________ Wikitech-l mailing list -- wikitech-l@lists.wikimedia.org To unsubscribe send an email to wikitech-l-leave@lists.wikimedia.org https://lists.wikimedia.org/postorius/lists/wikitech-l.lists.wikimedia.org/
wikitech-l@lists.wikimedia.org