Jump to content
Wikimedia Meta-Wiki

GLAM CSI/Final report

From Meta, a Wikimedia project coordination wiki
GLAM Wiki CSI - Contribution Study Initiative

GLAM Wiki CSI (Contribution Study Initiative) or GLAM CSI is a project to assess the contribution pipeline in the Wikimedia technical infrastructure for supporting cultural and heritage partnerships and projects. It ran from January to December 2024, and examined the ongoing efforts from the Smithsonian Institution and other cultural and heritage partners to produce real-life user stories from various GLAM-Wiki efforts.

The report was designed to document the needs and practices of content partners for participating in Wikimedia activities while identifying the opportunities for improvement and prototyping possible solutions, particularly on the shared focus of addressing knowledge gaps.

At the heart of the study was as survey of various stakeholders in the GLAM wiki community, to allow the research team to follow up with respondents to gain a deeper understanding of areas of interest.

The insights from this survey was used to:

  • Establish focus groups for more in-depth interviews
  • Develop first draft personas, user stories and user journeys
  • Identify tools that should be mentioned as helpful in the final report
  • Diagnose gaps in tools for GLAMs

User stories report

[edit ]

Background

[edit ]
Planned measure of success

(include numeric target, if applicable)

Actual result
Identify gaps in tools used by GLAMs to add or modify content to Wikimedia projects as an input for generating user stories The GLAM CSI project thoroughly reviewed existing tools (e.g., Pattypan, flickr2commons), highlighting specific functionalities that needed maintenance or updates. This allowed the team to craft user stories reflecting real-world tool limitations. Input was taken from online surveys and then refined in an extensive in-person session at Wikimania with the GLAM wiki community, creating a set of recommendations and additions to the listing at Wikidata:Linked open data.

Full report details can be found below. (Successful)

Learn about pain points faced by institutions in contributing to Wikimedia projects as an input for user journey creation Through interviews, surveys, and workshops with GLAM professionals, the project gathered insights into technical, legal, and organizational challenges. These findings were directly integrated into user journeys to illustrate key hurdles. A priority matrix of possible solutions, generated from a Wikimania 2024 gathering of GLAM professionals, was created with varying resource needs. (Successful)
Establish personas based on common trends observed in responses The project developed distinct contributor personas by analyzing recurring patterns in the data (e.g., archivist, curator, librarian). These personas helped tailor recommendations and resource development.

The draft user stories were first posted on meta for community feedback. Workshops at Wikimania and WikiConference North America refined them into the current set, ensuring they reflect diverse user types, varying challenge scales, and a broad range of knowledge domains. (Successful)

Prototype collaboration with DPLA to test a metadata-ingestion pipeline for Smithsonian units A joint proof-of-concept has been designed with DPLA's Dominic Byrd-McDevitt to map Smithsonian metadata to DPLA’s ingestion pipeline. DPLA briefly paused mass uploads in late 2024 which put some of the prototyping on hold, the partners continue the work into 2025 to enable ingestion and create a crosswalk database. (Ongoing, prototype deployment in progress)
Provide an analysis in supporting "advanced visual formats" in Commons for Smithsonian DPO (Digitization Program Office) 3D shaded/textured models To bring textured, full-color 3D models to Wikimedia Commons, we focused on adding support for the industry-standard glTF format. The need became urgent after the free, but commercially–controlled, Sketchfab service began winding down, leaving the 3D community in search of an open platform.

Our work revived stalled conversations, key Phabricator tasks (T187844, T246901), and mapped out the technical steps required for glTF integration. Activity in the Wikimedia 3D Telegram group and strong interest from GLAM partners eager to share high-quality models helped sustain momentum. (Partially successful; implementation under active discussion)

Survey

[edit ]

SURVEY: GLAM CSI survey for Wikimedia contributors

This survey was publicized in numerous platforms, email lists, community groups, and industry partners to examine the breadth of participation in Wikimedia activities, the challenges faced, and tools being used. Participants could respond on behalf of an institution, a specific project, a Wikimedia affiliate, or as an individual contributor.

Background

[edit ]

When Wikipedia was founded in 2001, engagement between the Wikimedia community and the cultural/heritage sector was not immediate. But each year saw Wikipedia's influence grow, as it showed up in Google searches, and was being linked to by people on blogs, social media, and even legal cases.

But a question always remained: How much could you trust Wikipedia's content and its volunteer community of contributors?

Wikimania 2008 at the Library of Alexandria, Egypt

Interest from respected cultural and heritage professionals gained traction in 2008, when the annual Wikimania conference was hosted by the Library of Alexandria in Egypt, signaling a shift in traditional institutions acknowledging Wikipedia's emergence as the world's foremost reference site.

In 2010, collaboration with the cultural sector took on the label "GLAM Wiki" when Liam Wyatt was the first Wikipedian-in-residence, at the British Museum in London. These initial collaborations took the form of joint public events such as edit-a-thons for article improvement, or the uploading to Wikimedia Commons of digital images of photos, documents, or collections with their associated metadata. Because of this increased need for contribution at scale, tools to mass upload image collections from GLAM institutions were developed to go beyond the standard Upload Wizard. These included specially developed utilities such as flickr2commons to allow "side loading" from existing license-compliant albums from users on flickr.com, or the development of commons:Commons:GLAMwiki Toolset Project by Europeana and EU-based Wikimedia chapters.

The advent of Wikidata in 2012 brought GLAM Wiki collaboration to a new level, particularly with libraries working with controlled vocabularies and data taxonomies. Wikidata's support for structured data and authority control records resonated with them, resulting in Wikidata training and sessions becoming a staple at top librarian conferences such as IFLA and LD4.

In 2016, the addition of Wikidata-like capabilities to Wikimedia Commons was initiated with the Structured Data on Commons project. This opened up more possibilities to full collections contributions by GLAM institutions, given the much wider scope of Commons, versus the stricter notability guidelines of Wikidata and Wikipedia.

However, in 2023 there was a significant concern about the uncertainty around the direction of GLAM Wiki efforts with a number of issues, many of which were highlighted by the Wikimedians in Residence Exchange Network User Group in a GLAM manifesto report and a session at Wikimania. Among the areas of concern were issues related to:

  • Contribution - Mass image uploading tools such as Pattypan were not being maintained, causing frustration for many GLAM partners
  • Enrichment - Uncertainty around the direction of Structured Data on Commons and the associated Commons Query Service
  • Measurement and evaluation - Many metrics tools became inoperable in 2023, leading to a number of GLAM institutions pausing their wiki efforts
Example of tools used by GLAM institutions in a Linked Open Data workflow. (From: wikidata:Wikidata:Linked open data workflow)

The report also highlighted progress: the Wikimedia movement has invested in OpenRefine, Thumbor and references for SDC during that period. However, there was overall concern about the general state of the tools as described in the Wikidata:Linked open data workflow.

It is within this climate that a proposal at the end of 2023 was made to better study and capture user stories and customer journeys to guide decisions going forward, with the Smithsonian Institution leading the study, in cooperation with the GLAM wiki community.

Funding

[edit ]

In 2024–25, the Smithsonian Institution received a Wikimedia Foundation grant to launch GLAM CSI, a project to examine the framework for traditional galleries, libraries, archives, and museums as partners, but also beyond that, to consider "complex contributions at scale," capturing detailed user stories, closing tooling gaps, and prototyping workflows that help Smithsonian units—and other GLAM partners—upload rich media and metadata to Wikimedia platforms.

As part of this initiative, the team surveyed and interviewed dozens of GLAM practitioners, distilled their insights into personas and user journeys, and prioritized critical pain points around batch uploads, structured data, and impact measurement. A key aspect of the prototyping effort is an enhanced ETLT (Extract, Transform, Load, Transform-again) ingestion pipeline, co-conceived with the Digital Public Library of America (DPLA). The project also investigates adding higher-fidelity media—especially textured 3D models—to Wikimedia Commons, addressing a long-standing need identified by the Smithsonian’s Digitization Program Office.

Expense Approved amount Actual funds spent Difference
Phase I: Pre-work
Project Management 10,744ドル.00 10,744ドル.00 0
Phase II: Implementation
Wiki Consultant Project Lead 90,000ドル.00 90,000ドル.00 0
Conference Registration & Travel 2,596ドル.00 2,458ドル.80 +137.2
Indirect 10,334ドル.00 10,320ドル.28 +13ドル.72
Total 113,674ドル.00 113,523ドル.08 +150ドル.92

Timeline

[edit ]

Planned milestones and feedback sessions at upcoming conferences and gatherings:

  • Q1 RESEARCH
    • Identify key stakeholders for interview and feedback, with a prioritization on the creation of user stories and customer journey map. Begin identification of key tools based on Content Partnerships Hub survey from 2022, extending it to new tools and facilities.
    • March - Finalize research survey in consultation with GLAM wiki community
  • Q2 GATHER AND PROTOTYPE
    • April - European GLAM coordinators online meeting describing survey and participation
    • Start production of user stories and journey map for feedback, gather detailed data around technical contribution pipeline. Begin pilot prototyping the use of DPLA /Smithsonian metadata and content uploading on a smaller Smithsonian units.
    • May 3-6 - Wikimedia Hackathon in Tallinn, Estonia, and GLAM + AI Sauna, Helsinki, Finland
  • Q3 EVALUATE
    • Publish drafts of intermediate results for feedback from key stakeholders and Wikimedia community. Incorporate feedback and prepare report.
    • mid-June - Intermediate results published and feedback sessions with GLAM community
    • August 7-10, 2024 - Wikimania 2024 in Katowice, Poland, as part of the GLAM track
      • Planned presentation of user stories results and evaluate feedback from GLAM wiki community
      • Brainstorming of solutions and demonstrating prototyping directions
  • Q4 REPORT
    • Refine report and begin drafting recommendations based on feedback. Begin dissemination and discussion of results at major gatherings including Wikimania, WikiConference, or high value gatherings (emphasis on November 2024 events).
    • October 3-6, 2024 - WikiConference North America in Indianapolis, Indiana, USA
      • Presentation of final user stories for refinement
      • Test prototypes of solutions and receive feedback on possible directions

Feedback on linked open data workflow tools

[edit ]

The project introduced the Linked Open Data workflow (Wikidata:Linked open data workflow) document to workshop attendees, which was new to many folks. After quickly explaining the various phases of this classic "Extract-Transform-Load" type workflow, we challenged the participants to respond to the chart with experiences and recommendations.

Participants were asked to individually provide feedback on each of these six phases of the workflow using sticky notes. People were given 20-30 minutes to walk amongst the posters in the room, noting to also read the stickies left by others. They were asked to leave notes in three areas for each column:

  1. Edits or additions to the list of tools
  2. Positive experiences
  3. Challenges

An extra poster to hold any "Other workflow ideas" was also provided.

Some rough overall observations could be made right away:

  • Ingestion (3) had the most varied and numerous responses, including many additions of specialized tools for ingestion, such as those related to iNaturalist, BHL, or SourceMD.
  • For ingestion solutions, there seem to be a good capacity for creating specialized or custom ingestion (3) tools, whether they were scripts, or custom code.
  • A number of feedback notes pertained to using tools to revise, adjust, or clean up data after an initial load of content - VisualFileChange.js, Quickstatements, or Cat-a-lot. This reflects comments that the ingestion/upload tools may not always do all that is desired by the user, requiring multiple tools to add all the relevant metadata, reflecting more of an iterative ELT (extract-load-transform) process rather than a traditional ETL process.
  • The most feedback was on the Reporting (6), with many concerns about the reliability and capability of measuring impact of contributions, both within the Wikimedia ecosystem, and especially for help measuring impact externally.

User stories

[edit ]
[edit ]

List

[edit ]

Overview

[edit ]

Distribution of user stories

Additional outcomes

[edit ]
At Wikimania, outcome of brainstorming, ranking, and sorting possible solutions to "obstacle areas." Summarized in chart below.

Prioritized solutions matrix

[edit ]

From the Wikimania Global Wiki Meetup – possible solutions to "obstacle areas" for Wikimedia work were generated, and ranked by participants from the in-person workshop. Together, these recommendations outline diverse paths forward—from technical solutions and better documentation to fostering a more inclusive, multilingual environment.

Hardest Harder Easier Easiest
GLAM council or consortium (6)

Language help: more tools for non-Roman alphabet (4)

An easy to handle surface for Wikidata queries (1)

Automated API refresh of GLAM-submitted data on Wikidata (2)

More technical support for Wikibase integrations (2)

Prioritize those tools that are easiest to use (3)

Impactful improvements to the platforms to make tooling easier (1)

Consistent branding (3)

Charge GLAMs for tech services (6)

Wiki community training in local context (4), truly multilingual (5)

Better onboarding, learning, training for partners to help them get started mapping data (2)

Layering/weighting data "official" vs contributed (5)

Start from project -> search tool (3)

Institutional "buddying up" for better understanding (4)

Training folks in tech documentation (6)

International GLAM events (3, 4)

Recognize low hanging fruit (tools and projects) (3)

Support tool developers when there are major changes (1)

Step by step recipes for institutions with entry level tasks, 'cookbook' (3) case studies (3)

Examples (case studies) of impactful reuse (5)

Ways to encourage reuse (5)

GLAM Wish List (6)

Systematic tools assessment system by quality (status) and importance (usefulness) (1)

Understand the future of Structured Data on Commons in the movement (2)

Legend: Languages and i18nTraining and documentationTools and development

Prototype project plan: DPLA ingestion pipeline for Smithsonian collections contribution

[edit ]

Introduction

[edit ]

How might we leverage the existing DPLA–Wikimedia Commons ingestion pipeline to support bulk uploads of Smithsonian image content, along with rich, high-quality metadata? DPLA has a proven track record of large-scale uploads for its partners, and the process can be further refined to meet the specific needs of various Smithsonian units.

Dominic Byrd-McDevitt (DPLA) and Andrew Lih (Smithsonian) propose an enhanced "ETLT" pipeline — Extract, Transform, Load, and Transform again — that incorporates a robust upload and post-upload metadata enrichment process. By embedding authority control identifiers at the time of upload (such as DPLA IDs, Smithsonian ARK IDs, and Smithsonian resource IDs), we can establish a streamlined, efficient upload process while enabling subsequent enrichment through additional processing passes, including scripts or bots that enhance integration with Structured Data on Commons.

Pilot project focus: NASM (National Air and Space Museum) CC0 collection

[edit ]

We selected the NASM ppen access collection—a compact, well-defined set of artifacts—to keep the project’s scope manageable.

Scope: ~1000 items

Consider an example NASM object from DPLA:

The object's DPLA metadata via API call:

This provides a JSON response with certain fields useful to DPLA:

Link to full object information from Smithsonian:

"isShownAt": "http://n2t.net/ark:/65665/nv9a33b50a4-3ff2-4f7c-81b3-ece5487393b7"

Individual media files for ingestion:

"mediaMaster": [

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_001-000001",

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_002-000001",

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_003-000001",

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_004-000001",

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_005-000001",

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_006-000001",

"https://ids.si.edu/ids/deliveryService?id=NASM-252A5F0C7CD72_007-000001"

],

After uploading files to Wikimedia Commons, DPLA will set some standard Wikibase properties for the Structured Data on Commons:

Property Value Derived from
Collection (P195) <Smithsonian unit>

Example: "National Air and Space Museum"

Q752669

DPLA JSON API:

"name": "National Air and Space Museum"

Inventory number (P217) <Smithsonian unit accession number>

Example: A20000640000

DPLA JSON API for some entities, such as Missouri Historical Society:

"identifier": [

"P0821-01-272"

]

Smithsonian entities seem to not be returning this field. We may need to deep dive into the XML field "originalRecord.stringValue" to extract this:

doc.descriptiveNonRepeating: <identifier label="Inventory Number">

NOTA BENE: This is especially challenging as this is not consistent across Smithsonian units! For example, the Smithsonian American Art Museum (SAAM) uses:

<identifier label="Object number">

Beyond this, we can request special fields to be set, as long as they can be derived from the standard API and XML fields.

Property Value Derived from
P7851 (Smithsonian resource ID) nasm_A20000640000 See P217 above. This can either be derived algorithmically, but it would be preferable to obtain this exact string from the API.
P9473 (Smithsonian ARK ID) nv9a33b50a4-3ff2-4f7c-81b3-ece5487393b7 DPLA JSON API, stripped from URL:

"isShownAt": "http://n2t.net/ark:/65665/nv9a33b50a4-3ff2-4f7c-81b3-ece5487393b7"

References:

Conclusions and next steps

[edit ]

The effort demonstrates (1) reliable extraction of source records by DPLA of SI metadata, (2) accurate mapping of metadata can be done within DPLA's existing pipeline and exposure to SI's API, and (3) automated post-upload enrichment that seeds Structured Data on Commons statements can be automated given key identifiers. Although edge-cases remain—such as inconsistent accession-number fields across Smithsonian units—the approach can work for for larger collections.

To build on this success, we recommend four next steps:

  1. Refine cross-unit mapping: Develop a shared "crosswalk" database that normalises variant field labels (e.g., Inventory Number vs Object Number) and stores unit-specific rules for identifier extraction.
  2. Automate post-upload enrichment: Package the enrichment scripts as a repeatable bot job, adding persistent identifiers (ARK, Smithsonian resource ID) and source references immediately after each upload batch.
  3. Increase ingestion frequency: Move from quarterly bulk dumps to a rolling or monthly schedule, giving curators and Wikimedians quicker visibility of newly digitised items and reducing backlog risk.
  4. Scale to additional Smithsonian units: After fine-tuning the workflow for NASM, onboard other Smithsonian museums (e.g., NMAH, NPG, NMAAHC) and invite external GLAM partners to test similar pipelines, establishing Commons as a sustainable, open alternative to commercial platforms.

By following this roadmap, DPLA and the Smithsonian can turn a one-off prototype into a standing pipeline—bringing high-quality, well-structured cultural heritage content to the Wikimedia ecosystem at scale.

Advanced visual formats: enhancing 3D content support on Wikimedia Commons

[edit ]

Introduction

[edit ]

Wikimedia Commons serves as a central repository for free-to-use media files, supporting the Wikimedia movement's mission of sharing free knowledge. While Commons has embraced various multimedia formats, its support for 3D content has been limited to basic STL files which began in 2018. STL, however, is a basic format geared towards object printing and does not support colors or textures, which restricts the quality and usability of 3D models. There is a need for supporting more modern file formats, with requests coming directly from GLAM partners and Wikimedia community with knowledge preservation and interactive applications in mind.

2. Motivation for Improvement

There is a growing need to support advanced 3D formats to enhance the quality and educational value of 3D content. Collaborations with GLAM institutions, such as the Smithsonian's Digitization Program Office, have highlighted the potential of 3D models for academic and cultural preservation purposes, with more than 3,000 high-quality textured open access models ready to contribute. Additionally, within the GLAM movement, groups such as Wiki World Heritage User Group and WikiKsour seek to use more advanced photogrammetry and 3D modeling of cultural sites, which can only be supported with more advanced formats like glTF.

A significant industry development is also a factor – Sketchfab is a commercial platform owned by the company behind Unreal Engine and Fortnite. While it initially supported open access and community sharing through a generous free tier, its acquisition by Epic Games in 2021 and subsequent business shifts have led to the elimination of that tier. This has become especially problematic for museums, educators, and Wikimedia contributors who have relied on Sketchfab to share culturally significant 3D content using formats like glTF. Its policy change threatens the accessibility and preservation of thousands of openly licensed models, highlighting the urgent need for sustainable, nonprofit alternatives such as Wikimedia Commons to support the public interest. In December 2025, Smithsonian's 3D digitization lead, Vince Rossi, approached the Wikimedian at Large Andrew Lih and the Commons community to see if we might expedite support for glTF. This led to the revival of the Wikimedia 3D Telegram group and active discussion of a prototype on Wikimedia Cloud.

In short, the recent changes in Sketchfab's operations, which will remove the free tier, have created an urgent need for a new platform where the 3D community can share their content freely.

3. The Case for glTF

The Graphics Library Transmission Format (glTF) is a modern 3D file format that supports textures, materials, and animations, offering a richer and more immersive 3D experience. The Wikimedia community and GLAM partners have shown strong support for adopting glTF to improve the quality and utility of 3D content. This shift would not only enhance visual fidelity but also attract more contributors and users, fostering a richer repository of 3D educational resources.

4. Challenges and Implementation

Supporting new file formats like glTF involves technical challenges, including software integration, interface adjustments, and community training. File formats are not often added to Wikimedia communities. Therefore, the tools and procedures for incorporating new file formats, both technically and culturally, are usually complex.

However, the good news is that support for more multimedia and 3D content was recently the most actively voted-on item in the Community Wishlist, showing a high level of interest.

5. Recommendations and Next Steps

To successfully integrate glTF support, we recommend:

  • Timeline and Resources: Establishing a clear timeline and allocating necessary resources for development and testing. Phabricator tasks:
  • Collaboration: Strengthening partnerships with GLAM institutions and the 3D content community to ensure a smooth transition and adoption, and understanding what kind of user interfaces would be needed.
  • Long-Term Vision: Envisioning a future where Wikimedia Commons becomes a leading platform for diverse and high-quality 3D content, thereby enriching free knowledge resources globally.

AltStyle によって変換されたページ (->オリジナル) /