For prior discussion on this talk page, see /Archive 1.

Works vs Editions

edit

WikiProject:Books follows FRBR, which is a system used internationally be libraries, to distinguish between data recorded for a work and data recorded for an edition. On Wikisource, the two should be placed on separate data items. This way, multiple editions of a work can each have their data stored on Wikidata, without confusion. Each published edition should be separate from all other editions, and separate from a data item for the work of literature. --EncycloPetey (talk) 16:10, 30 March 2023 (UTC)Reply

For example: Death Comes for the Archbishop (Q115978630) is the 2023 Standard Ebooks edition. And I have created Death comes for the Archbishop (Q117344844) for the 1927 edition being transcribed at en.WS. These two editions must have their own data items, or else there would be multiple dates, publishers, distribution formats, and links all jumbled together on the same data item. --EncycloPetey (talk) 16:42, 30 March 2023 (UTC)Reply

@encycloPetey: Thank you. So https://librivox.org/death-comes-for-the-archbishop-willa-cather would have its own item, or would it be whichever edition they are reading from? — user:Arlo Barnes (talk) 08:14, 31 March 2023 (UTC)Reply
If it's possible to identify which edition they read from, then it could be placed on that edition. However, that's usually difficult or impossible to determine. Additionally, we need to track information like the reader, publish date for the recording, the cast members if it's a LibriVox perfromed play, or the readers of each section if it's a divided effort reading. So the best option is to give LibriVox recordings their own data items. --EncycloPetey (talk) 16:14, 31 March 2023 (UTC)Reply

VLM

edit

Are you going to start a discussion on creating a VLM property, so we can import the entire database and do a mix-n-match? RAN (talk) 19:36, 11 December 2023 (UTC)Reply

@Richard Arthur Norton (1958- ): Yes, please comment at wikidata:property proposal/Veterans Legacy Memorial. Arlo Barnes (talk) 19:37, 11 December 2023 (UTC)Reply
And now, Veterans Legacy Memorial ID (P12389)! 06:17, 1 February 2024 (UTC)

Land acknowledgement question

edit

Hi@Arlo Barnes! I am reaching out to gather some ideas about the approach of adding land acknowledgement information to Wikidata items.

I am working with a group of people on a wikidata project. We wanted to add land acknowledgement to some institutions. Our intuition was to search if there's property to use. We then found that Wikidata users add the land acknowledgement by adding "subject use (P2283) land acknowledgement (Q96200400)" statement and add statement URI as a reference. Some users use "of" qualifier to add information about the groups mentioned in the land acknowledgement. It seems that "of" is being deprecated (https://www.wikidata.org/wiki/Property:P642). We wonder if you know what will happen to the land acknowledgement statements after "of" is deprecated?

We also see if there's property for other kind of statements. There's also another property for accessibility statement URI https://www.wikidata.org/wiki/Property:P9494. We discussed if we can propose a land acknowledgement URI property, which seems more convent for users to add this information. We wonder if you have concerns of creating a new property for land acknowledgement URI?

Our group also found there was a property proposal land acknowledgement deleted recently (https://www.wikidata.org/w/index.php?title=Wikidata:Property_proposal/land_acknowledgement&action=edit, @Ameisenigel) We wonder if you have thoughts on this topic as well?

Thanks! Gretaheng18 (talk) 17:48, 12 July 2024 (UTC)Reply

'of' has been labelled as deprecated for years, I don't think anyone is going to remove it until there's a clear solution for transitioning the (many!) statements to another property. Proposing a dedicated property for such acknowledgements seems like a good idea to me and I'd vote for it, although I'll note that I haven't seen most institutions use a webpage for that; it looks like they usually just put it in the footer of the site, or alongside legal notices. In any case, there's always an opportunity to provide the exact phrasing used with quotation or excerpt (P7081) as a qualifier or quotation (P1683) in the references.
Although not directly related to land acknowledgement (Q96200400)s, you may also be interested in meta:Wikiproject Local Contexts. Arlo Barnes (talk) 18:26, 12 July 2024 (UTC)Reply
Just noting that Wikidata:Property proposal/land acknowledgement was containing only the template for property proposals but did not contain any information. Feel free to request a new property, but please fill out all required fields. --Ameisenigel (talk) 20:53, 12 July 2024 (UTC)Reply

Invitation to participate in research exploring Wikidata editors' experience of content gaps

edit

Dear Arlo_Barnes,

I hope this message finds you well!

We are researchers at King's College London investigating how content gaps arise and can be measured in Wikidata. Currently we have explored existing research papers to identify several categories of gaps. However, we have noted a lack of consideration of editors’ experiences in this existing research and are keen to hear about editors’ views on and methods for identifying and addressing content gaps. While this topic has seen a lot of attention in Wikipedia, we believe Wikidata presents unique challenges and content which warrant further investigation.

We are reaching out as we understand you have previously taken part in a research study with our colleague Kholoud and thought as an active and experienced editor, you may be able to share your experiences with us. We would therefore like to invite you to participate in an interactive online workshop to explore this topic further.

This will consist of a 90 minute online call consisting of a group discussion and collaborative editing of a document (in Miro). We will ask you to rank gaps according to your familiarity and your opinion of their importance, give feedback on which types of metric might be most valuable to you as an editor and give some initial thoughts and potentially even sketches of how a content gap monitoring tool might look for Wikidata.

The main goal of the workshop is to understand your perspectives on how to measure and monitor content gaps as well as potentially identify further metrics or even to propose new methods to identify and quantify gaps in Wikidata.

Participation is completely voluntary. All personal data will be kept confidential in compliance with GDPR. If you are interest in taking part, you can find out more about the workshop from our participant information sheet. You can also read more about the research at our meta page.

You can sign-up to take part from our registration form.

The workshop will take place online using Microsoft Teams. We are hoping to host the workshop in the coming weeks, but if you would like to take part and are unavailable during the proposed times, we may be able to find an alternative time.

If you have any questions or concerns, please do not hesitate to contact us at anelia.kurteva@kcl.ac.uk and neal.t.reeves@kcl.ac.uk

Thank you for considering taking part in this workshop and supporting our research.

With kind regards, Celestialtoast (talk) 20:25, 28 September 2024 (UTC)Reply

Dear Arlo_Barnes,
Thank you so much for kindly volunteering to participate in our study. Unfortunately, it was difficult to find a time when everyone was available and we ultimately had to arrange the workshop for 4pm UK time tomorrow which the sign-up form suggests would be a time you weren't available. We're hoping to hold another workshop soon to give further insight into our findings and if it is ok with you, I would be happy to reach out again when we've made the arrangements.
Just in case and on the chance that you are able to attend after all and it's not too short notice, I wanted to share the session link.
Thank you again for your willingness to take part and sorry that we couldn't find a more convenient time.
Celestialtoast (talk) 20:47, 8 October 2024 (UTC)Reply

Appetite of a People-Pleaser (Q125087673)

edit

Hi. audio track (Q7302866) should only be used for non-music audio tracks, not songs Trade (talk) 04:04, 8 October 2024 (UTC)Reply

Thanks for developing the concept of land acknowledgement

edit

Bluerasberry (talk) 22:26, 6 January 2025 (UTC)Reply

Wikidata:Property proposal

edit

There is a lot of property proposals that needs votes (opposing or supporting) in order to move along. Would you mind helping out? Trade (talk) 02:20, 2 February 2025 (UTC)Reply

I'll take another look later today; skimming the topic groupings right now I see a lot I am ambivalent about, but I want to be able to comment with something more substantive than "support because properties shouldn't be a pain to make" or "oppose because we're kinda neglecting the properties we already have"... Arlo Barnes (talk) 18:45, 2 February 2025 (UTC)Reply

Just a thought

edit

Have you ever considered requesting to have your own bot on Wikidata? There's plenty of tasks i can think of that would need it but most of them seem to be too busy to take on any additional maintenance tasks Trade (talk) 05:12, 17 March 2025 (UTC)Reply

Yes, I've often wished for the capability for narrow automation. For example, talk:Q136208. However, I am not quite where I wish to be in terms of understanding of operating bots, nor having the local setup to run one. Help in the former would be appreciated, if you have knowledge in this area. Arlo Barnes (talk) 01:45, 19 March 2025 (UTC)Reply

P642 use cases

edit

I would be interested in what you find confusing about the use cases page. Obviously it's large (that's unavoidable), but for any given statement there should only be a few tables that may be relevant, which can be found from the ToC. Obviously, I don't think your diagram adds much value, certainly not enough to justify its prime real estate, but maybe there are other things we can do to make the page more usable. Swpb (talk) 13:28, 19 March 2025 (UTC)Reply

I freely admit the diagram is currently a dud, but I think that's the sort of thing I personally would need to navigate that page. Perhaps my cognition is unusual in that respect, but in order to figure out (to use my latest use of P642 as an example) which statement(s) retire all government employees (Q133355473) should have to replace main subject (P921)dismissal (Q9213592)P642public employee (Q3796928), I have to determine the animacy of...government employees? Or dismissal? Or should I just give up on modelling that nuance? What I need is a clear signal of where the 'intuitive' but imprecise modelling that P642 provides may fall short of the precision required by Wikidata. I wish for the traceable paths of a flowchart, even if a literal flowchart is perhaps the wrong tool for the job here. Arlo Barnes (talk) 00:30, 20 March 2025 (UTC)Reply
Thanks for the reply. So, I think of the ToC as providing essentially a flowchart already, and I want to see if I can understand why it doesn't work that way for you. Is it just hard to trace visually? Because I certainly don't like what the new Vector skin has done with the ToC, but you can disable that for yourself by clicking "switch to old look", so at least the ToC will have numbering, better indentation, and no line wrapping. I wish I could force that display for everyone.
Now, as to your specific example. Government employees are certainly animate (being living humans), but that doesn't matter here, because they are not the cause of their own dismissal, but its undergoer/patient. (This is where I think the new ToC layout could be tripping you up, with the line break in "agent (animate cause)" making the word "cause" easy to miss.) So, knowing the qual value is an undergoer, we are looking at this table. There, you won't find a case covering the occurrence "dismissal of employees", until you get to the "catch-all" case at the bottom of the table. Here I admit it gets a little tricky – I wanted a single property covering all three of these relations, but I could only get enough support for this three-way split. Anyway, "government employee" is a role filled by a person, rather than something a person inherently is; in a different context, the same person may have roles like "mother", "customer", or "passenger". So we want the last property, objects of occurrence have role (P12992). Using class of object(s) of occurrence (P12913) would technically be wrong, but it's a subtle point and not the end of the world. As a hint, groups of people will almost always demand the "role" property, and I've added a note to that effect.
One more thing to note is that the final arbiter of whether a property is correct is the property item itself: its description, examples, and constraints. The use cases page is a best effort at directing you to candidate properties, but you don't have to rely on it to decide between those candidates. That's why step 3 is "make sure the handling makes sense". Cheers!
Edit: One more thing: as to "where the 'intuitive' but imprecise modelling that P642 provides may fall short of the precision required by Wikidata", the answer is "everywhere". That's why it's going away. Swpb (talk) 14:20, 20 March 2025 (UTC)Reply

Help

edit

Could you help me complete HDMI 2.0 connector (Q133806867) and DisplayPort 1.4a connector (Q133806868) and power connector (Q133806881)? Trade (talk) 05:29, 3 April 2025 (UTC)Reply

I can take a look tomorrow. What's missing? Arlo Barnes (talk) 06:08, 3 April 2025 (UTC)Reply
Nvm. It's already fixed
If you care me and PantheraLeo1359531 are trying to figure out how to model graphic cards Trade (talk) 11:17, 3 April 2025 (UTC)Reply

Interest in collaborating with Decolonise Wiki

edit

Hi @Arlo Barnes: — nice to meet you!

I’m Victor Murari, a Brazil-based art historian working on digital colonialism in contemporary art. I came across the Decolonise Wiki group and it aligns closely with my current research and two projects I’ve been developing: Decolonial Atlas (mapping artistic practices that confront digital colonialism) and Interrupted Narratives (an experimental VR exhibition about algorithmic control and curatorial meta-platforms).

I’d love to learn how to get involved (open tasks, norms, meetings, preferred channels). Here is a brief one-pager of my work: [1]. Happy to share full portfolios if useful.

Is there a preferred onboarding pathway for newcomers to Decolonise Wiki?

Thanks so much! — Victor Victorvtm (talk) 12:28, 15 October 2025 (UTC)Reply

Usually a WikiProject's talk page would be a good place to make an intro, but since you already messaged everyone involved, no need. I think the project is still in early stages, so any input you have could help guide it. It seems there are already items for some of the artists you cover in the Decolonial Atlas:
Not yet Juan Pablo Pacheco Bejarano, Davida Enara, Mayara Ferrão, Fanuel Leul, Pule Magopa, Gabriel Massan, oZzo UKumari, Yinkore, or Kadu Xukuru, so if there are handy sources about them to help fill out some items according to notability guidelines, that could be a next step. I see you posted a blog post about Wikidata, so I presume you are already familiar, but if you would like help I'd be happy to join in. Arlo Barnes (talk) 06:39, 16 October 2025 (UTC)Reply
@Arlo Barnes: Many thanks for this thorough roundup and the concrete pointers. I’m currently in conversation with the other two people listed in the project, and there is interest in continuing to work on this together. I fully agree with your suggested next step and will start gathering solid sources for the artists who don’t yet have items (e.g., Juan Pablo Pacheco Bejarano, Davida Enara, Mayara Ferrão, Fanuel Leul, Pule Magopa, Gabriel Massan, oZzo UKumari, Yinkore, Kadu Xukuru).
I'm still developing my skills and knowledge on the platform, but I expect to feel more comfortable very soon and to contribute more actively. I’ll report back here as I make progress so we can coordinate.
Thanks again for the guidance! — Victor (Victorvtm) --~~~~
Victorvtm (talk) 11:55, 16 October 2025 (UTC)Reply

Wikidata Platform Newsletter - March 2026

edit


This is the 4th issue of our monthly newsletter! The next issue will be published in April 2026.

  • Team communication update: Since November, our small and newly formed team has been developing our approach to sharing our work with the community and incorporating feedback. Migrating Wikidata Query Service’s infrastructure is a large and complex undertaking that serves many different audiences. Our goal has been to find the right mix of channels, engagement levels, and response times that allows us to reach the broadest possible audience with the resources we have.
Now that we have more of this structure in place, we are sharing project updates through regular newsletters and publishing reports and learnings on-wiki, where discussion pages are open for questions and comments. Our Phabricator board is available for flagging bugs or engaging in ongoing tasks. We also host regular office hours as a shared space for community members to raise questions and topics that may be relevant to others as well (next is in April). Questions added to the Etherpad will be addressed during those sessions.
The team reviews feedback shared through these channels and incorporates it where it helps advance our migration goals. While we aim to respond to questions within about a week when possible, we may not be able to reply to every individual point.
  • Label and MWApi service migration : We’ve completed an analysis of the wikibase:label and wikibase:mwapi services. One or both of these services are used in more than half of the requests sent to the Wikidata main graph. Preserving the functionality of these features is a core requirement of our backend migration away from Blazegraph. Full results of this investigation can be found on Wikitech.
  • Rate limits: Global rate limits for Wikimedia Foundation APIs were announced on March 2 and will be rolled out over the next month. These limits do not currently apply to WDQS, as we are still evaluating appropriate rate limits for our platform as part of our backend migration (ETA: July ‘26). In the interim, to protect against service interruptions like the ones observed during the week of February 23 following an increase in request volume and complexity, we have rate-limited a handful of identified users whose requests gridlocked our system. While details regarding specific actors have not been publicly disclosed (PII is involved) Wikimedia SRE deployed a workaround to a known Blazegraph bug that manifests under traffic spikes (T242453). While this does not solve the root cause of the problem, we expect it to increase WDQS reliability in the short term. We will continue to implement these spot fixes as needed until a more scalable solution is determined.
  • Work in progress: traffic analytics and operations: In the February sprint, we dedicated time to improving analytics on traffic and query behavior, with a particular focus on query latency classes and volumes by user-agent. This work provides two main benefits: 1) it informs different access patterns and latency-bound quality-of-service considerations, and 2) it led to short-term improvements in the real-time metrics we use to operate WDQS. We introduced new panels in Grafana, as well as internal-facing analytics tools, to report error rates and latency buckets, as well as new alerts that trigger on trends indicating timeouts. This approach allows us, on one hand, to proactively react to traffic spikes before an outage impacts end users, and on the other hand, to define the observability requirements that our new target backend will need to meet.
  • Work in progress: development and test infrastructure: In our previous newsletter, we published exploratory benchmarking that identified two candidate systems for Blazegraph replacement that met minimum requirements. The purpose was to gain experience with open source triple store implementations as we embark on the migration. Building on these learnings, one of our goals this quarter is to deploy test instances of Virtuoso and QLever on internal eqiad infrastructure, to gain experience operating both systems and supporting internal development efforts (T414443). We have now 4 hosts that serve main and scholarly data via QLever and Virtuoso (respectively). This work informed improvement ideas for our data infrastructure, and allowed us to refine our understanding of indexing times on graph splits. As part of this work, we modified the real-time index updater code base to remove hard dependencies on Blazegraph. While this is a work in progress, we can now update both QLever and Virtuoso indexes in real-time while reusing existing WDQS infrastructure (T414447). This allows us to experiment and load test traffic on infrastructure similar to Blazegraph-based WDQS, and is a milestone towards exposing a new backend to public traffic later this year. If you have experience operating Wikidata triple stores at scale, we would like to get in touch and exchange learnings.


Udehb-WMF (talk) 17:07, 12 March 2026 (UTC)Reply

Introducing PPM: A Tool to Streamline Wikidata Property Proposals

edit

Hi Arlo Barnes, I wanted to share a tool I’ve been working on: Help:Property Proposal Manager (I upgraded since you installed). It’s a comprehensive script designed to streamline the whole property proposal workflow on Wikidata.

It covers creating proposals, voting, closing discussions, following up, marking unsigned comments, handling maintenance and auto-fixes, and even supports one-click property creation. The idea is to make managing the full lifecycle of property proposals a lot smoother and less time-consuming.

I’m not entirely sure how useful it might be for you, but I thought it could be worth a look. I’d really appreciate any feedback or suggestions if you get a chance to try it. Regards, ZI Jony (Talk) 03:32, 9 April 2026 (UTC)Reply

Thanks for the heads-up. I rarely make proposals, but the next time I do I'll definitely use this tool. Arlo Barnes (talk) 20:08, 9 April 2026 (UTC)Reply

Wikidata Platform Newsletter - April 2026

edit


This is the 5th issue of our monthly newsletter! The next issue will be published in May 2026.

  • Q3 Wrap-up: We closed out the month of March by completing all of our team goals for the quarter. This work included setting up test environments for production-replay traffic testing using on-premises hardware, the next step from last quarter’s benchmarking analyses, defining data access guidelines for our platform, and solidifying details of our migration plan such as release schedule and target use cases. The results of our Q3 efforts are currently being reviewed with internal stakeholders and will be shared with the community soon as part of our broader communication on the upcoming Blazegraph migration.
  • Q4 Plans: In Q4, the Wikidata Platform team is shifting from planning to execution readiness. We are focused on finalizing our migration plan, building out automated query validation to ensure correctness of the new service, developing a communications plan to smooth the transition, and beginning the technical build of our new architecture. In parallel, we are continuing work on platform access and Quality of Service (QoS) guidelines, including rate limiting and user authentication policies aimed at reducing system abuse, improving telemetry, and aligning with broader WMF standards. We will share more details and updates throughout the quarter.
  • Upcoming Feedback Cycles: Within this month of April, we will share artifacts outlining our recommendation for a Blazegraph replacement and our proposed technical architecture for the Wikidata platform. We welcome community feedback on our plans-your input will help us smooth the transition and develop mitigations for impacted use cases. We will share more details about the upcoming migration once we’ve collected and incorporated community feedback, including a timeline of changes, query rewriting best practices, and migration support documentation for the soon-to-be launched endpoints.
  • Team is Growing: Our team is growing! At the beginning of March, we welcomed Andrea Westerinen to the team as a contractor. In the coming months, Andrea will assist with the upcoming Blazegraph migration by creating technical documentation, providing query rewrite support, and advising on SPARQL requirements.
  • Rate Limiting Updates: As announced in last month’s update, we are working closely with site reliability engineers (SREs) to protect against WDQS service interruptions caused by increased request volume and complexity. While we work towards migrating to a more scalable backend that’s more resilient than Blazegraph, we are selectively rate-limiting users when we observe a correlation between their activity and performance incidents. Additionally, we have deployed and fine-tuned auto-remediation measures to restart servers that have been gridlocked. The intention of these changes is to stabilize the WDQS experience for the community, which we have observed to be effective. If you believe you’ve been negatively impacted by these changes, please let us know.
  • Runbooks for Common WDQS Issues:
    • Data reconciliation: Some WDQS users reported an issue where items deleted from Wikidata were still queryable (T407702). Here’s some background to explain the issue: WDQS is continually synced against Wikidata via a streaming-update data flow with automatic retrying in case of failure, but it can happen that an update is missed and never gets picked up by the streaming updater. This can cause Blazegraph’s state to drift from Wikidata’s. In this case, we determined that the Wikidata deletion events were indeed missed by the streaming updater, and we resolved the issue by manually running a tool that rereads the source of truth for a specific entity (or list of entities) and updates Blazegraph accordingly. We also added a Wikitech runbook to document this fix. After the migration we will have more options available for keeping WDQS’s data up to date, including regular bulk data syncing and reindexing pipelines.
    • High Lag troubleshooting: Over the last two quarters the Wikidata Platform team has been working with SRE teams to troubleshoot and proactively address queries and actors that are putting a strain on WDQS infrastructure. We reviewed existing metrics, alerts and Grafana panels, and tuned our instrumentation to emerging traffic patterns. We authored and shared a new runbook that guides Wikidata Platform engineers in troubleshooting and escalation, during periods of elevated WDQS load and user-facing query failures. This work is part of a broader operational excellence effort in which we are partnering with SRE and other teams at the Foundation to streamline cross-team collaboration and protect our infrastructure.
  • Wikidata through Wikimedia Enterprise: On March 31st, Wikimedia Enterprise announced the launch of Wikidata APIs as part of the suite of Enterprise offerings. These endpoints are built to support high-volume access needs by commercial reusers. This new resource will provide right-sized solutions for some of the largest consumers of Wikidata and users of our platform, supporting our vision for stable and sustainable access to Wikidata for everyone. More details can be found in WME’s full announcement.


Udehb-WMF (talk) 17:20, 13 April 2026 (UTC)Reply

Eshanebul

edit

Hi Arlo,

I think you may be the person I'm looking for.

I've been trying to trace the origin of "Eshanebul" — the Láadan name for Istanbul that appears on Wikidata. The edit history points to you, 26 August 2025, 04:21.

The word seems to have traveled from that single entry into nine small Wikipedia editions (Tatar, Lombard, Ga, Twi, Sylheti, Karakalpak, Talysh, Tyap) via bots — the only appearances of the word anywhere on the internet. I found it tonight by accident, tracing a rabbit hole that started on the Gagauz Wikipedia.

I'm curious : did you coin it yourself, or were you recording a word that already existed somewhere in the Láadan community? And does it follow Láadan word-building rules — is there an etymology to it ?

I've genuinely been at this for hours. You're the end of the trail. — Arigâteau18 (talk) 11:04, 25 April 2026 (UTC)Reply

Hi, this is a pleasant surprise; I had no idea that Wikipedias would import the name due to it being the only name (P2561) statement on that item. If that is causing problems I'd like to help figure out how to fix them.
But to answer your questions, the word didn't exist previously, and generally most vocab items have yet to be established for Láadan. Unlike common nouns and other parts of speech that would pragmatically require a community process to develop, proper nouns can be generated by adapting the sounds to Láadan's phonology (although of course the best thing would be if the name came out of the place in question).
For awhile there was Lexemes Challenge (Q109617338), which suggested topics to add WD:LEXicographical data for. I thought that was why I added it, but now I can't find which challenge would have involved Istanbul. edit: b:en:template:Hello, Jonathan! may also have had something to do with it...
The name itself doesn't have etymological parts, just phonetic ones; since Istanbul means 'to the city' one could gloss it as 'Miwithedim' (from miwith (L1563772) and -dim), but I didn't. There are no consonant clusters in Láadan, nor is there an unmodified 's' sound, so Ist- becomes Esh-, and -anbul becomes -anebul.
See also: names of European cities in different languages (I–L) (Q17083748)Arlo Barnes (talk) 14:39, 25 April 2026 (UTC)Reply

Wikidata Platform Newsletter - May 2026

edit


This is the 6th issue of our monthly newsletter! The next issue will be published in May 2026.

  • Backend replacement and architectural design proposals: For the past few months, we've shared learnings from our investigations into backend replacement, and today we are sharing our recommendations for a new RDF database to replace and accompanying technical architecture for the migration away from Blazegraph as the backend of the Wikidata Query Service (WDQS). These documents outline the selected direction and how the new architecture is designed to improve scalability while making future backend changes easier.
We are inviting feedback from the Wikidata community and other WDQS users until 25th May 2026, particularly on:
  1. Any important considerations we may have missed
  2. How your WDQS use cases, tools, or workflows may be affected
We encourage you to share feedback on the migration discussion page. You can also join our upcoming office hour (Next tomorrow, May 12th) to ask questions and discuss your use cases.
To help identify higher-risk areas, we have created a page to track high-impact use cases and tools. This page is not intended to catalogue all WDQS usage, but to highlight complex or critical cases that may require additional attention.
In preparation, we encourage you to add questions, feedback, or migration-related support needs to this etherpad. This helps us shape the agenda and focus on the most relevant topics during the session.


Udehb-WMF (talk) 14:34, 11 May 2026 (UTC)Reply

Possible ontology course project?

edit

Hi Arlo! It was nice to meet you on Tuesday, and good to virtually see you again today. I was really happy to see you joined WP:P244 maintenance. For the course project, I wonder if you would be interested in working with me on making subclasses to organize the Help:Deprecation/List of reasons for deprecation. I think IagoQnsi has the right idea with this list (and appreciate that you commented there), but I wonder if projects like WP:Data round-tripping and WP:Authority control might benefit from being able to query for all deprecated statements whose errors fall under a class named something like... well, I don't know really. "Reasons for deprecation that suggest a problem with the source data"? And then things like conflation (Q14946528), duplicate entry (Q1263068), and error in referenced source or sources (Q29998666) would be an instance of that class? And we'd have other kinds of classes for other kinds of deprecation reasons?

I'm still thinking this through, as I'm sure you can tell. But I wanted to reach out to you because if you're interested in working on this together, we should probably start talking this through together sooner rather than later. If this doesn't interest you at all (which is fair! This is very much coming out of what I would find useful for my work) then I won't worry too much about setting up a proposal for this project this weekend.

Anyways, please let me know what you think of this very amorphous, undefined idea. Looking forward to chatting with you!

(Also, sorry to be nosy. I saw your note on IagoQnsi's page, then clicked on your contribs. Installing removeredunantlabels.js might be more convenient for removing labels in favor of mul than working by hand. See here for more gadgets. The one Peter was using in class is classification.js) Mcampany (talk) 19:38, 14 May 2026 (UTC)Reply

It's okay to be nosy as far as my contributions are concerned, although I'm terribly inconsistent with editing rigour. That said, thanks for the link to the 'remove redundant labels' script, I've installed it now.
One particular area of deprecation I've been interested in but unsure if/how to proceed is values of HTTP status code (Q110861089) for qualifying statements of official website (P856) and related properties. HTTP 404 (Q404) seems to get used a fair amount but 451 Unavailable For Legal Reasons (Q17055519) not so much, for example. I'm not sure if that's the kind of thing you mean since it isn't really a type of error in the source, but in getting at a source in the first place. Arlo Barnes (talk) 20:04, 14 May 2026 (UTC)Reply
Hmm... I wonder if those might be better defined as something like "Reason for deprecation because of an access error"? Because as you mentioned, those are all issues with accessing a website, rather than an issue with the data itself. It may be helpful to have a category for things like HTTP 404 (Q404) though. I'm not sure if Wikidata has infrastructure like the Internet Archive bot running around like English Wikipedia does. If there isn't, having a category of deprecations for link rot (Q1193907) may help to surface that issue.
Anyways, I'll start a project page tomorrow or over the weekend, and I'll drop a note here when it's up. No pressure if you decide you'd rather do another project--this is a pretty niche use case, and I've been tempted by some of the other offerings on the project page myself. ;) Mcampany (talk) 21:10, 14 May 2026 (UTC)Reply
Hi Arlo! I hope you had a relaxing weekend. I've put together a project proposal page. Feel free to add yourself if you're interested, and if it's not appealing, no worries! :) Mcampany (talk) 18:35, 18 May 2026 (UTC)Reply

Strewn field vs strewn-field-producing meteorite

edit

Hi, the item strewn field (Q1059884) was about the concept of strewn field. On 2020-08-24 you changed the English label and description to say that it's about a meteorite that results in a strewn field. Why do this instead of creating a new item for the latter? QIDs should keep referring to the same concept. --Tengwar (talk) 22:59, 27 May 2026 (UTC)Reply

You're right, it's simple error. I'll separate the concepts. Arlo Barnes (talk) 02:25, 28 May 2026 (UTC)Reply
Thanks! --Tengwar (talk) 22:25, 28 May 2026 (UTC)Reply

Wikidata Platform Newsletter - June 2026

edit


This is the 7th issue of our monthly newsletter! The next issue will be published in July 2026.

  • QLever as the New Backend System for WDQS: After reviewing feedback from the community on our architectural proposals shared last month (see Backend Replacement and WDQS Architecture Re-Design), we have decided to migrate to QLever as the new backend for the Wikidata Query Service. This decision has been aligned upon across partners in WMF, WMDE, and the QLever team. We are very excited about the improvements to performance and sustainability that this choice unlocks for WDQS users. We will continue to share details on the user-facing impacts of this decision, in future newsletters and on our migration project page, as the work progresses.
We encourage all members of the community to continue helping us identify higher-risk areas by reporting use cases on our high-impact use cases and tools page. This page is not intended to catalogue all WDQS usage, but to highlight complex or critical cases that may require additional attention.
  • Migration Timeline: Following the decision to use QLever, we would like to share some key milestones for our migration. More details will be shared as we approach full implementation. Please note that all dates are targets and may change in the event of unforeseen challenges. Refer to the migration project page for up to date information.
    1. Exploration: (COMPLETED) From September through March, the team conducted traffic and benchmarking analyses to understand the needs of WDQS users and alternatives to our current system. This culminated in our final recommendations for a new backend and platform architecture, which have been reviewed and aligned upon across stakeholders.
    2. Installation: (IN-PROGRESS) In April, the team began building. The development of new QLever endpoints (ie. WDQS v2) is underway, as is the refactoring of our platform architecture. This includes work on indexing, update pipelines, and rewriting observed production traffic into the SPARQL 1.1 standard. We are on track to complete our build of the new endpoints by July 1st and transition into initial implementation. The query service will continue to be available in its current state through implementation phases.
    3. Initial Implementation: (NOT STARTED) WDQS v2, a new endpoint built on QLever, will be first available to a small number of pilot users. The team will work closely with this group to learn where improvements are needed and how we can best support users in independently migrating their use cases. A self-service hub will be published by October 1st. This will include learnings from our pilot group, guidance on how all users of WDQS can migrate their work flows, and documentation on best practices for adapting all Blazegraph dependencies to our new system or alternative endpoints where needed.
    4. Full Implementation: (NOT STARTED) The new endpoints will be scaled to meet the needs of the broader community and will be generally accessible to all. The original Blazegraph endpoints will still be available, but may experience service degradation beginning in February, 2027 as we reallocate resources to the new infrastructure and begin slowly winding down the legacy service. We aim to have all WDQS traffic migrated by June 30, 2027, at which time the Blazegraph endpoint will be decommissioned.
  • Query Categorization and Testing: We have begun evaluating WDQS queries to identify Blazegraph-specific features and functionality. All bespoke query aspects need to be rewritten into the SPARQL 1.1 standard in order to work with the new QLever backend. We have documented our process for this work in two Wikitech publications on SPARQL Query Characterization and Test Architecture for QLever. These documents provide more detail on methodology and testing used to validate the correctness and performance across the new and old backends. They are intended as supporting technical references for contributors interested in the migration validation and benchmarking approach.
  • 2026-05-08 incident report: On May 7, 2026, aggressive web scrapers began overwhelming the Wikidata Query Service (WDQS), triggering a multi-day service degradation that impacted both availability and lag SLOs. The excessive load caused Blazegraph to timeout for over half of users at peak, while also throttling the streaming updater, which blocked index updates and cascaded into edit throttling on wikidata.org itself. Initial mitigation on May 7-8 included depooling the eqiad datacenter WDQS deployment and applying rate limits based on sampled request data, but the outage persisted through the weekend. Full resolution came on Monday, May 11, when deeper analysis of WDQS logs revealed a scraper that had evaded sampled webrequest data used for initial rate limiting. Once a targeted rate limiting rule was applied to the scraper's signatures, timeout rates returned to normal. See the full incident report
  • Blazegraph Migration Office Hour (June session): Our next Blazegraph Migration Office Hour will take place on Tuesday, 9 June 2026 (Tomorrow) at . This session is focused on supporting the migration away from Blazegraph as the backend of WDQS. Whether you have questions, need clarification, or want to discuss how your use case may be affected. You can register for the session via the event page.
In preparation, we encourage you to add questions, feedback, or migration-related support needs to this etherpad. This helps us shape the agenda and focus on the most relevant topics during the session.


Udehb-WMF (talk) 11:10, 8 June 2026 (UTC)Reply

Wikidata Platform Newsletter - July 2026

edit


This is the 8th issue of our monthly newsletter! The next issue will be published in August 2026.

  • Documentation on Query Rewrites: As WDQS continues its migration off Blazegraph, queries that rely on Blazegraph-specific extensions will need to be updated to comply with the SPARQL 1.1 and GeoSPARQL standards. To support this, we are publishing documentation to help you identify affected queries and rewrite them, so your queries keep working through the migration and beyond.
This transition to standard SPARQL reduces vendor lock-in, improves interoperability with other triplestores, and ensures our codebase remains robust. Importantly, the specifics of these rewrites are doing more than just guiding manual code updates. The details define the processing logic for an upcoming tool designed to automatically convert Blazegraph queries into standard SPARQL for the QLever engine.
The query rewrite details are focused on four key areas:
  1. Graph Analytic Services (GAS): See Blazegraph_Migration:_Rewrite_of_GAS
  2. MediaWiki API (MWAPI) Service: See Blazegraph_Migration:_Rewrite_of_MWAPI
  3. Label and Utility Services: See Blazegraph_Migration:_Rewrite_of_Label_and_Utility_Services_and_Functions
  4. Geospatial Functions: See Blazegraph_Migration:_Rewrite_of_Geospatial_Services_and_Functions
We Want Your Feedback!
As we map out these transition rules and develop the automated QLever conversion program, your input is critical. Please review the Wikitech pages linked above and let us know whether:
  • The rewrite rules cover your use cases
  • There are edge cases or specific Blazegraph quirks that are not accounted for
Please share your thoughts, concerns, or examples of queries on the related Discussion pages of the links above, so we can ensure the automated converter works seamlessly for everyone.
The system architecture from the WDQS v2 design doc has been implemented and deployed on Kubernetes. We ran a successful end-to end-test and validated that the QLever deployment can pick up a newly built index, start to backfill it with real-time events, and once ready the database can be queried. The service returns valid results, metrics are properly reported and available in Grafana, and logs are shipped to Wikimedia’s Observability Platform. Upcoming work will be focused on performance optimization and improvements to the indexing and real-time update pipelines supporting the service.
  • Timeline Update: As mentioned in our June newsletter, we are kicking off the Initial Implementation phase of the Blazegraph migration in July. After a few short weeks of final testing and iteration on the new endpoints, the first round of WDQS v2 users will be given access to the new endpoints so they can begin their migration. The QLever endpoints will become accessible to the broader community by October 1st, when we kick off Full Implementation, along with documentation on query rewriting.
  • Reminder to Report Use Cases: We encourage all members of the community to continue helping us identify higher-risk areas, namely SPARQL queries that are highly complex, by reporting use cases on our high-impact use cases and tools page. As a reminder, any use case can be reported through this mechanism. The purpose of this page is to give the WDP team visibility into user-level migration needs.
  • Blazegraph Migration Office Hour (July session): Our next Blazegraph Migration Office Hour will take place, Tuesday, 7 July 2026. This session is focused on supporting the migration away from Blazegraph as the backend of WDQS. Whether you have questions, need clarification, or want to discuss how your use case may be affected. You can register for the session via the event page.
In preparation, we encourage you to add questions, feedback, or migration-related support needs to this etherpad. This helps us shape the agenda and focus on the most relevant topics during the session.


Udehb-WMF (talk) 10:54, 3 July 2026 (UTC)Reply

Wikidata Platform Newsletter - August 2026

edit


This is the 9th issue of our monthly newsletter! The next issue will be published in September 2026. This is the 9th issue of our monthly newsletter! The next issue will be published in September 2026.

  • WDQSv2 Scaling: Work in Progress: The new QLever public endpoints are serving live traffic and are generally available to pilot cohort clients for testing. Our partner team at WMDE started porting the Query UI to https://query-next.wikidata.org and https://query-scholarly-next.wikidata.org. We have also rolled out a number of improvements to support federation and increase availability of the service. We have stood up a new staging environment on Kubernetes. Lastly, we are experimenting with a bleeding edge version of QLever that significantly improves memory footprint and update performance, and supports index rebuilds without needing service restarts.
  • SPARQL Query Re-Writing Tool: We have begun work on a tool that will rewrite Blazegraph-specific SPARQL queries into a format that is compatible with our new QLever implementation. This tool will apply the rewrite guidance we have published on WikiTech, offering an easier path for users to migrate to WDQS v2. The purpose of this tool is not to fully adapt all queries to the new endpoint, but to support users through their migration efforts. Development on the test infrastructure is underway and expected to be complete by August 14th. An evaluation of rewrite performance will be complete by September 4th, when we will share our results. We will iterate throughout the month of September and provide access to the tool by the first week of October.
  • Reminder to Report Use Cases: We encourage all members of the community to continue helping us identify higher-risk areas, namely SPARQL queries that are highly complex, by reporting use cases on our high-impact use cases and tools page. As a reminder, any use case can be reported through this mechanism. The purpose of this page is to give the WDP team visibility into user-level migration needs.
  • Blazegraph Migration Office Hour (August session): Our next Blazegraph Migration Office Hour will take place tomorrow, Tuesday, 4th August 2026. This session is focused on supporting the migration away from Blazegraph as the backend of WDQS. Whether you have questions, need clarification, or want to discuss how your use case may be affected. You can register for the session via the event page.
In preparation, we encourage you to add questions, feedback, or migration-related support needs to this etherpad. This helps us shape the agenda and focus on the most relevant topics during the session.


Udehb-WMF (talk) 17:10, 3 August 2026 (UTC)Reply