[{"content":"Hello! I work as a postdoctoral researcher in the ACES team at Télécom Paris and am primarily interested in team dynamics in open source projects. This blog post gives a bit of context around my current research focus.\nI am also involved in some open source projects myself. Currently:\nForgejo, a forge platform Mergiraf, a conflict resolution tool for Git Debian, a popular Linux distribution Previously:\nOpenRefine, an open source data cleaning tool EditGroups, a tool to undo faulty batches of edits on wikis Dissemin, a platform to encourage researchers to publish their articles in open access and many others that you can find on my Codeberg and GitHub profiles.\nI am also mildly active in other spaces:\nthe W3C Entity Reconciliation Community Group, where we develop a protocol for online data matching the CAPSH association, doing advocacy around open access to scientific publications I am based around Paris, France.\nYou can also find me on:\nMastodon: @pintoch@mamot.fr Matrix: @pintoch:matrix.org Codeberg and GitHub Contact me by email: [my first name]@[my last name].eu Also, check out my band!\n","date":"26 September 2026","permalink":"https://antonin.delpeuch.eu/about/","section":"FOSSilitation","summary":"\u003cp\u003eHello! I work as a postdoctoral researcher in the \u003ca href=\"https://aces.telecom-paris.fr/\" target=\"_blank\" rel=\"noreferrer\"\u003eACES team\u003c/a\u003e at \u003ca href=\"https://www.telecom-paris.fr/\" target=\"_blank\" rel=\"noreferrer\"\u003eTélécom Paris\u003c/a\u003e\nand am primarily interested in team dynamics in open source projects.\n\u003ca href=\"/posts/improving-organization-management-in-forgejo/\"\u003eThis blog post\u003c/a\u003e gives a bit of context around my current research focus.\u003c/p\u003e\n\u003cp\u003eI am also involved in some open source projects myself. Currently:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"https://forgejo.org/\" target=\"_blank\" rel=\"noreferrer\"\u003eForgejo\u003c/a\u003e, a forge platform\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"https://mergiraf.org/\" target=\"_blank\" rel=\"noreferrer\"\u003eMergiraf\u003c/a\u003e, a conflict resolution tool for Git\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"https://www.debian.org/\" target=\"_blank\" rel=\"noreferrer\"\u003eDebian\u003c/a\u003e, a popular Linux distribution\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003ePreviously:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"https://openrefine.org\" target=\"_blank\" rel=\"noreferrer\"\u003eOpenRefine\u003c/a\u003e, an open source data cleaning tool\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"https://github.com/Wikidata/editgroups\" target=\"_blank\" rel=\"noreferrer\"\u003eEditGroups\u003c/a\u003e, a tool to undo faulty batches of edits on wikis\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"https://association.dissem.in/dissemin-closure.html\" target=\"_blank\" rel=\"noreferrer\"\u003eDissemin\u003c/a\u003e, a platform to encourage researchers to publish their articles in open access\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eand many others that you can find on my \u003ca href=\"https://codeberg.org/wetneb\" target=\"_blank\" rel=\"noreferrer\"\u003eCodeberg\u003c/a\u003e and \u003ca href=\"https://github.com/wetneb\" target=\"_blank\" rel=\"noreferrer\"\u003eGitHub\u003c/a\u003e profiles.\u003c/p\u003e","title":"About me"},{"content":"FOSSilitation is a blog about the Free and Open Source Software (FOSS) movement. I write about the daily failures of collaboration that I witness and share my facilitation experiments.\nThere is an RSS feed for it.\n","date":null,"permalink":"https://antonin.delpeuch.eu/","section":"FOSSilitation","summary":"\u003cp\u003eFOSSilitation is a blog about the Free and Open Source Software (FOSS) movement.\nI write about the daily failures of collaboration that I witness and share my facilitation experiments.\u003c/p\u003e\n\u003cp\u003eThere is an \u003ca href=\"posts/index.xml\"\u003eRSS feed\u003c/a\u003e for it.\u003c/p\u003e","title":"FOSSilitation"},{"content":"","date":null,"permalink":"https://antonin.delpeuch.eu/posts/","section":"Posts","summary":"","title":"Posts"},{"content":"Remember Sourceforge, where you\u0026rsquo;d download software in the 2000s? It\u0026rsquo;s now mostly dead, and the company running it has been trying to squeeze out as much profit from it as they can using questionable practices.\nBut despite that, it\u0026rsquo;s a software forge, and I think some aspects of its design remain interesting and could even be an inspiration for today\u0026rsquo;s forges. Let\u0026rsquo;s have a look at a sample FOSS project hosted on Sourceforge, Code::Blocks, with which I made my first programming steps some decades ago.\nCompare that to your favorite GitHub project. There are so many differences! Here\u0026rsquo;s what I find striking.\nA landing page designed for end users #The Download button jumps to your face. There\u0026rsquo;s also a list of reviews (more about this below) and a \u0026ldquo;Support\u0026rdquo; Tab. The platform assumes you\u0026rsquo;re here to install, use, and get help for this software.\nIt\u0026rsquo;s clearly a user-centric design.\nCompare that to the experience on GitHub. If you want to download a tool hosted on GitHub, you\u0026rsquo;ll need to, either:\nknow your way around things and choose the \u0026ldquo;Releases\u0026rdquo; tab, take the latest release, scroll past the release notes down to the artifacts, select the one that\u0026rsquo;s relevant for your operating system based on the filenames (which can be somewhat cryptic) and download that or scroll past the file listing to the README, make your way through a forest of CI and code coverage badges, to reach a download link hopefully added by the developers to the README, probably still leading you to the release page anyway, where you\u0026rsquo;ll need to continue with the above! It\u0026rsquo;s mind-boggingly hostile!\nGitHub has made the deliberate choice to tailor the landing page of a project to a developer-centric viewpoint. Which probably explains its success among developers.\nBut how can developers afford to ostracize their users to that extent? Is it because software distribution is happening through other means anyway, such as app stores? Is it because they\u0026rsquo;re expected to make a website for their software (for instance via GitHub Pages) and keep users there? But then is it honest to expect users to file issues on GitHub when something doesn\u0026rsquo;t work as expected? Or take part to GitHub Discussions? Or even just leave a GitHub star? As an end user, you\u0026rsquo;re not even invited to create an account on GitHub anyway: it\u0026rsquo;s meant for developers, period. I can\u0026rsquo;t help but think that it would be preferrable not to drive such a wedge between the users and developers of a tool.\nMore meaningful than stars: reviews #You can actually give your opinion about a piece of software on Sourceforge! You can even rate specific dimensions: \u0026ldquo;Ease\u0026rdquo; (of use), \u0026ldquo;Features\u0026rdquo;, \u0026ldquo;Design\u0026rdquo; and \u0026ldquo;Support\u0026rdquo;. And you can write a comment on top of that! Really cool!\nIn comparison, GitHub stars are used as a proxy for project success all the time, but it\u0026rsquo;s a really poor signal. As a user I only have a binary choice: do I give a star to this project or not? And even if I do, the meaning is unclear: I can star a project because I\u0026rsquo;m a happy user of it, or because I want to remember to have a look at it later, or because I find the concept behind the repo funny…\nReadable description comes before file listing #Notice how the landing page doesn\u0026rsquo;t show a listing of the files in the repository? Instead, it shows a human-readable description of the project. The file listing is still accessible via the \u0026ldquo;Files\u0026rdquo; tab.\nThis makes so much sense, and not just for users, but also for developers. As a developer, when I go to a GitHub repository, a significant area of the page is just dedicated to showing me the file listing of the root directory, which will generally match the standard directory structure of whatever programming language or framework is in use. Of course, as a developper I can be interested to know which technology is used for this project, but this is a super inefficient way to convey that information.\nHere\u0026rsquo;s my internal discourse when landing on such a page (slightly exaggerated, I concede):\nOh, they have a .github directory! They must be using CI with GitHub actions. And maybe Dependabot too. That\u0026rsquo;s good. And then they have a src directory as well. Good! I love software which has… source code in its source code repository. And look what else is coming! A tests directory! Wow, they are really rocking it, they got a test suite. And next… a coveralls.yml! Ah, the good old Coveralls service! I wonder if they know about CodeCov. A package-lock.json! This smells like a Javascript application, interesting :) …\nYou got the idea: I don\u0026rsquo;t need to see all those files when I\u0026rsquo;m just discovering a project.\nBrought to you by a team, not an irreplaceable founder #Notice how the header of the page says:\nBrought to you by: killerbot, mandrav, mortenmacfly, thomas-denk\nHow refreshing! There\u0026rsquo;s a team behind this project!\nIf we were on GitHub, it would rather look like this, with one of the maintainers listed as single owner and the others added as \u0026ldquo;collaborators\u0026rdquo; (whose list is not publicly visible):\nor, perhaps somewhat better, the repo would be owned by a \u0026ldquo;CodeBlocks\u0026rdquo; organization where all maintainers would have been added, with only two of them being publicly visible because the others didn\u0026rsquo;t take the initiative to publish their membership… sigh.\nI\u0026rsquo;m convinced that GitHub\u0026rsquo;s design has serious consequences on the way people relate to the projects they are involved in. By designating the repository creator as owner, the platform is making it harder to meaningfully share responsibilities, leading to burnout or abandoned software. I wrote earlier about the implicit feudalism that this creates.\nAdvertising related projects #Another interesting feature of this landing page is the side-bar with recommended similar projects. The first one it lists is CBFortran, \u0026ldquo;A customized distribution of the Code::Blocks IDE for Fortran\u0026rdquo;.\nIt might not seem like much, but I find it an amazing feature.\nI wrote earlier about the poor discoverability of forks on GitHub. The problem is simple: if a software project becomes unmaintained, it is very hard for the community around that software to coordinate and agree on a new repository where to continue the maintenance. All they can do is open an issue in the original project, where they announce the fork, but that will not be visible to people who arrive on the landing page of the project. If they are not able to discover an actively maintained fork, that means they might end up starting their own. This means a forest of forks, a ton of duplication of work, and a lot of frustration.\nI don\u0026rsquo;t actually know how Sourceforge\u0026rsquo;s recommendations are computed and whether an active fork would organically appear there, or if it needs to be somehow vetted by the project maintainers. But I find it encouraging to see an alternative distribution for Fortran being listed in first position.\nAny other feature? #You made it to the end of my love letter to Sourceforge. Congratulations. As mentioned in previous blog posts, I\u0026rsquo;m currently working on Forgejo and those design considerations are motivated by that work. So stay tuned as I turn Forgejo into the new Sourceforge, with ads everywhere, promoted links and partner deals! ;)\nIf you have other pointers to interesting features, or have any comments about the thoughts above, I\u0026rsquo;d love to read them on Mastodon.\n","date":"18 September 2026","permalink":"https://antonin.delpeuch.eu/posts/the-forgotten-features-of-sourceforge/","section":"Posts","summary":"\u003cp\u003eRemember \u003ca href=\"https://sourceforge.net/\" target=\"_blank\" rel=\"noreferrer\"\u003eSourceforge\u003c/a\u003e, where you\u0026rsquo;d download software in the 2000s?\nIt\u0026rsquo;s now mostly dead, and the company running it has been trying to squeeze out as much profit from it as they can using questionable practices.\u003c/p\u003e\n\u003cp\u003eBut despite that, it\u0026rsquo;s a software forge, and I think some aspects of its design remain interesting\nand could even be an inspiration for today\u0026rsquo;s forges.\nLet\u0026rsquo;s have a look at a sample FOSS project hosted on Sourceforge, \u003ca href=\"https://sourceforge.net/projects/codeblocks/\" target=\"_blank\" rel=\"noreferrer\"\u003eCode::Blocks\u003c/a\u003e, with which I made my first programming steps some decades ago.\u003c/p\u003e","title":"The forgotten features of Sourceforge"},{"content":"At this year\u0026rsquo;s Wikimania (the annual conference of the Wikimedia movement), a new sub-event was introduced: the Team Challenges. This post summarizes my experience of it as a participant.\nTraditionally, the Wikimania hosts a rather loosely structured \u0026ldquo;hackathon\u0026rdquo;, which mostly consists in making a big room available for people who want to hack on things throughout the conference. There is an opening ceremony for it, where participants are able to pitch the projects they intend to work on, and a closing ceremony where they present what they did.\nThis works well for folks who know what to work on and how. It is also a generally quiet place where I can get a quick break from more social activities happening in the rest of the conference. But it\u0026rsquo;s clearly not for newcomers: you need to know what to do. The people around you are all staring at their screens and doing their things: if you don\u0026rsquo;t know them yet, it\u0026rsquo;s not so easy to break the ice or ask for help.\nThe Team Challenges offered a very welcome alternative to this format. We signed up to participate a few months ahead of the conference, stating our prior experience, areas of interest and languages spoken. Based on those sign-ups, the organizers assigned us to teams, where we had different roles:\nparticipant, not necessarily familiar with the Wikimedia tech ecosystem but wanting to get involved mentor, with prior experience of building things for Wikimedia projects, available to support their team members to build something together super-mentor, someone with deep knowledge of a specific domain, who can be reached out to by anyone in the event (not just from their team) coach: the organizers invited coaches from EMCC (or rather its French branch) to facilitate the work in teams In my team, there were one coach, one super-mentor, one mentor (myself) and five other participants. A few of them I knew to various degrees from various online and in-person interactions. More infos about the structure of the event can be found on the event website. I\u0026rsquo;ll focus on my own impressions and lessons learned.\nHelping newcomers get started with tool development for Wikimedia #This format felt very appropriate to help people get started in the Wikimedia tech spheres. The focus was on team work, not on just building a prototype as fast as possible. The participants were coming from more diverse horizons than what we generally experience at the hackathon.\nOf course, there is a big question mark as to what the participants do with this experience after the end of the event: are they going to continue this sort of work? Do they have the confidence, time and knowledge to do so outside of the supportive structure of the event? I expect it to be rarely the case, but that\u0026rsquo;s true of most events organized around the onboarding of new members in a community: there is necessarily a significant proportion of drop-outs. Also, even if they don\u0026rsquo;t stay, they might still have got something out of it that can be useful in other contexts.\nHelping old timers get better at mentoring #When I signed up for the event, I did so somewhat reluctantly because the format of the event was a bit daunting. By signing up as a mentor, I was hoping I\u0026rsquo;ll just be asked some questions from time to time, such as how to deploy a tool as a gadget or on toolforge, or help people navigate the different possibilities to consume data from Wikidata.\nIn the first phase of the challenges, happening online, I was quickly thrown into what felt like a much different role: scheduling and facilitating our first online meeting and trying to find a process to identify a common project for us to work on. It felt quite a bit more managerial that I was expecting, and also required more time than I expected ahead of the event. But it was definitely good training!\nWorking with a coach #When the organizers told the mentors they invited some coaches to participate, there were mixed reactions in the advisory group around the event. Someone mentioned having had bad experiences with coaches. Involving coaches who had no prior exposure to the Wikimedia movement looked perhaps a bit risky.\nIn our team, it felt like a positive experience to me. The presence of our coach Aurélie felt really useful to me, especially in the first online phase where we were all a bit lost as to what to work on and how to get to know each other. The separation between my mentor role and her coach role was a bit unclear to me, because from the organizers' communication I had the impression that I was expected to take the lead and facilitate the team-building more than her, but I asked for her help with those initial meetings and it worked out well enough, I would say.\nHaving someone on the team with minimal prior exposure to Wikimedia meant that we sometimes had to make some things more explicit and quickly explain some terms, but that wasn\u0026rsquo;t a big hindrance and probably helped create a climate where it was okay to ask about basic things.\nSo I\u0026rsquo;m really glad they took this initiative.\nCommunication with the organizers #The communication ahead of the beginning of the online phase was somewhat sparse and I was overall a bit confused about what sort of work we were expected to do before the in-person event. I was a bit stressed about our team being able to deliver something by the end. Because this was a brand new format, it was particularly challenging for the organizers since we couldn\u0026rsquo;t rely on prior experience.\nIt also felt a bit difficult to have to have to adapt to changes in the team composition, which happened multiple times after the start of the first phase. It was time-consuming and perhaps somewhat detrimental to the team spirit to have to onboard latecomers after the fact.\nBut those were minor annoyances - I am really happy that the organizers tried out this new format, which was a success in my opinion. I\u0026rsquo;m also quite proud of our team and had a great time. I hope to see this format in other events in the future, be it within Wikimedia or in other contexts.\nI\u0026rsquo;ll leave you with this sketch of our group by David Revoy, who was part of the team! We are represented in the style of the \u0026ldquo;Zéfirottes\u0026rdquo; creatures, magical creatures invented by Claude Ponti, after which our team was named.\nCC BY 1.0 David Revoy\n","date":"29 August 2026","permalink":"https://antonin.delpeuch.eu/posts/wikimania-2026-team-challenges-experience-report/","section":"Posts","summary":"\u003cp\u003eAt this year\u0026rsquo;s \u003ca href=\"https://wikimania.wikimedia.org/wiki/2026:Wikimania\" target=\"_blank\" rel=\"noreferrer\"\u003eWikimania\u003c/a\u003e (the annual conference of the Wikimedia movement),\na new sub-event was introduced: the \u003ca href=\"https://wikimania.wikimedia.org/wiki/2026:Team_challenges\" target=\"_blank\" rel=\"noreferrer\"\u003eTeam Challenges\u003c/a\u003e.\nThis post summarizes my experience of it as a participant.\u003c/p\u003e\n\u003cp\u003eTraditionally, the Wikimania hosts a rather loosely structured \u003ca href=\"https://wikimania.wikimedia.org/wiki/2026:Hackathon\" target=\"_blank\" rel=\"noreferrer\"\u003e\u0026ldquo;hackathon\u0026rdquo;\u003c/a\u003e, which mostly consists in making\na big room available for people who want to hack on things throughout the conference. There is an opening\nceremony for it, where participants are able to pitch the projects they intend to work on, and a closing ceremony\nwhere they present what they did.\u003c/p\u003e","title":"Wikimania 2026 Team Challenges: experience report"},{"content":"When you want to clone a large git repository locally, consider using:\ngit clone --filter blob:none \u0026lt;the_remote_repository\u0026gt; This is so much better than --depth 1! Instead, this method still gives you the full commit history but only downloads the file contents on demand, as they are needed. Initially, only the files appearing in the last revision of the default branch are downloaded (to make the first checkout).\nYou can checkout any other revision and Git will download the necessary files from the origin on demand. Or run any other Git command, like git diff between commits you haven\u0026rsquo;t checked out yet.\nMind blown. Thanks to Théo Zimmermann for pointing me to that! See the documentation of this --filter option for more details.\nP.S.: running git blame with this set up is not optimal, as versions of the file are fetched one by one. But it still works! Maybe there is some optimization to do in this case.\n","date":"26 May 2026","permalink":"https://antonin.delpeuch.eu/posts/cloning-large-repositories/","section":"Posts","summary":"\u003cp\u003eWhen you want to clone a large git repository locally, consider using:\u003c/p\u003e\n\u003cpre tabindex=\"0\"\u003e\u003ccode\u003egit clone --filter blob:none \u0026lt;the_remote_repository\u0026gt;\n\u003c/code\u003e\u003c/pre\u003e\u003cp\u003eThis is so much better than \u003ccode\u003e--depth 1\u003c/code\u003e! Instead, this method still gives you the full commit history but only downloads the file contents on demand, as they are needed.\nInitially, only the files appearing in the last revision of the default branch are downloaded (to make the first checkout).\u003c/p\u003e\n\u003cp\u003eYou can checkout any other revision and Git will download the necessary files from the origin on demand. Or run any other Git command, like \u003ccode\u003egit diff\u003c/code\u003e between commits you haven\u0026rsquo;t checked out yet.\u003c/p\u003e","title":"Cloning large repositories"},{"content":"In January, I started a 2-year research project on the governance of FOSS projects. I do this as part of the ACES team at Télécom Paris, where I am hosted by Théo Zimmermann and Stefano Zacchiroli. The project is part of the CONGRATS area of the eNSEMBLE funding programme, which looks at the dynamics of online collaboration more broadly.\nOne part of my project consists in looking into the existing governance dynamics in those projects, trying to identify patterns that work better than others. This involves collecting data from large collections of FOSS projects (via the Software Heritage archive), interviewing people, and analyzing those results. I\u0026rsquo;ll write more about this later.\nIn this blog post I want to focus on another line of work I have been following in parallel: designing and implementing governance-related improvements in Forgejo, a forge platform used by a growing share of the FOSS ecosystem.\nWhy are forge platforms relevant to the governance of the projects they host? #Forge platforms (such as GitHub, GitLab or Forgejo) are popular solutions to coordinate the development of open source software. Beyond hosting files, they allow communication (through tickets or code review discussions) and permissions management, letting teams define various roles and assigning rights to them. The Continuous Integration (CI) systems often associated to forges also allows to define advanced rules and procedures tailored to the project.\nThrough my experiences in various FOSS projects, I have become convinced that the design of a forge platform has deep implications for the governance of the projects it hosts. In small projects in particular, the participants rarely dedicate much time to negotiate explicit agreements about how they want to function as a team. The consequence of that is simple: the rules of collaboration boil down to what the forge platform allows.\nThis state of affairs, which can also persist even in not-so-small projects, can lead to many issues. Some of them are well-documented, such as the Ruby Central crisis which started when maintainers of Ruby-related repositories were removed unilaterally without notice. The abandonment of the nvim-treesitter repository which I wrote about earlier is another example: a maintainer, upset by a comment made by a user, decides to unilaterally archive the repository. Because they can. They do so without coordinating with other contributors who might have been happy to continue the work. Beyond such crisis incidents, I believe that the design of the forge platform has many subtle implications on the participant\u0026rsquo;s mental representation of their rights and responsibilities in the project, how they are credited for their work, how permissions are granted and revoked, and probably much more. I find Nathan Schneider\u0026rsquo;s thoughts on the \u0026ldquo;implicit feudalism\u0026rdquo; that those platforms create particularly convincing on this matter.\nHow can we improve the governance of FOSS projects at large? #I see two ways:\nto encourage more projects to think about their governance and formalize it, distancing themselves from the default \u0026ldquo;who\u0026rsquo;s gonna stop me\u0026rdquo; model offered by the forge; to improve forge platforms so that the default governance they imply is healthier, for instance by preventing uncoordinated destructive actions by design. Both of those are very ambitious. For the first approach, it\u0026rsquo;s unrealistic to expect that FOSS contributors would be able to dedicate significant efforts to maintaining governance structures in tiny projects: the effort needs to be proportionate to the size of the community and not overshadow the development and maintenance of the software itself. Beyond that, this first approach is about changing the culture of an entire movement, which cannot happen overnight. But if we switch to the second approach, it\u0026rsquo;s also unrealistic to expect forges to implement governance models fully. Seth Frey has an excellent blog post explaining why it might not even be desirable: What if governance technologies are the last thing we need?\nDespite those difficulties, I think it\u0026rsquo;s still worth working in both of those complementary directions. On top of that, there is a third way I am interested in:\nto improve forge platforms so that they ease the use of explicit governance models in the projects they host. With this third way, it\u0026rsquo;s not about expecting the forge platform to implement all of the governance mechanisms that a project could want to adopt. Instead, I only want the forge to provide a sensible substrate of governance features, making it easier for teams to follow the governance they have adopted. I \u0026ldquo;just\u0026rdquo; want the forge not to stand in the way of the agreements people have reached. The actual \u0026ldquo;implementation\u0026rdquo; of the governance would remain mostly manual, possibly augmented by external tooling (such as CI configuration or bots) when appropriate.\nWhy work on Forgejo? #Forgejo is a relatively young and agile open source project, with a project team that seems aligned with those goals. There is a steady stream of projects migrating to Forgejo, be it on Codeberg or by hosting their own Forgejo instance, such as Fedora or FFMPEG. I also have prior experience as a Forgejo contributor and have a deep appreciation for quite a few amazing team members.\nWhat improvements am I looking at? #Before considering more ambitious improvements, I have been working on small fixes in the domain of member management in Forgejo organizations, which also serves as a warming-up exercise.\nFor instance, the workflow to add a new member to an organization was quite convoluted, so I designed and implemented a better UI for it. This improvement is already deployed on Codeberg. There were also issues with organizations with lots of members, as some lists weren\u0026rsquo;t properly paginated and user avatars were served at a high resolution, leading to long loading times and poor usability.\nI am now working on design proposals for more ambitious features, and seek the feedback from the community on those. Here is an overview of the areas I have identified. I don\u0026rsquo;t expect to work on all of them: I want to see which proposals meet the most enthusiasm and go for those. I\u0026rsquo;m also open to shifting my attention to other improvements if they are deemed more promising.\nProvenance tracking for memberships #When viewing the list of members of an organization or team, I am not given any information about how those people got their membership. Were they added by someone else? Did they create the organization? Were they added because they authenticated via OAuth from a certain provider? I am thinking of adding this information to the members list, similarly to what GitLab does.\nForgejo links: issue #12567\nVisibility of team memberships #Many FOSS projects try to work in the open as much as possible and expect that the composition of their teams is publicly visible. Currently, as an outsider to a Forgejo organization, I can only see the members who explicitly chose to publicize their membership, and I cannot see which exact role they have in the project.\nI am considering introducing new settings at the organization and team level enabling to publicize those memberships systematically. To make sure that members consent, I am working on making it possible to decline joining a team, instead of being added directly as is currently the case. In my opinion, this opportunity to decline an invitation would make it acceptable to enforce the visibility of all team members (as they have all consented to that by joining it).\nForgejo links: issue #862, design proposal\nGit-based management of organization settings #Some projects store the list of their members (and their roles) in a Git repository. This list is then kept in sync either manually (for instance in xmonad or Mergiraf) or using some automation (in Guix or in Kubernetes). Doing so has the benefit of adding some transparency and offering a natural place to discuss membership changes as pull requests. That is particularly useful given the lack of native workflows for applying to a team, nominating someone for a position and discussing it, proposing the offboarding of inactive members, and so on.\nForgejo could potentially make it possible to create such a \u0026ldquo;magic repository\u0026rdquo;, whose contents would be kept in sync with the members of the organization. This could potentially be extended to cover other organization settings. Such a feature would avoid the need for setting up external tools (such as Peribolos, used in Kubernetes) as those require using API tokens with administrative rights and may run into rate-limiting issues, for instance.\nForgejo link: design proposal\nDiscussion space for membership changes #An alternative to the above would be to offer a native UI in the forge to propose and discuss membership changes. Think of them as discussions similar to pull requests, but associated to an organization instead of a repository, and covering a membership change (such as onboarding or offboarding from a team, which could potentially be proposed by anyone). Let\u0026rsquo;s call them \u0026ldquo;nominations\u0026rdquo;.\nThis would cater for a simpler UX, which wouldn\u0026rsquo;t require familiarity with a particular serialization format to represent lists of members in a git repository. It would also likely be safer and simpler to implement, to the cost of some flexibility (for instance, a given \u0026ldquo;nomination\u0026rdquo; would only be able to make changes to the membership status of one person at a time). It would likely be harder to extend the feature to cover other organization settings or actions, although there is interest for something along those lines.\nForgejo link: design proposal\nContacting users privately #While most conversations in FOSS projects should happen in public, there are cases where private communication is a better option. This is for instance the case when disclosing vulnerabilities, attempting to defuse a conflict, or coordinating an in-person meeting, for instance. Sarah Novotny makes this argument much better than I can in her blog post Open source needs private spaces.\nIt is not uncommon for forge users to carefully hide their identity, sometimes without leaving any way of contacting them. Forge platforms like Forgejo or GitHub don\u0026rsquo;t have any private messaging features (one convoluted way would be to create a private repository, invite the recipient to it, and then create an issue in that private repository).\nInstead of implementing yet another private message system in a forge like Forgejo, one could take inspiration from the way MediaWiki offers this feature. This platforms offers an \u0026ldquo;Email this user\u0026rdquo; form, which lets any logged-in user send an email to another user through this form. This doesn\u0026rsquo;t disclose the email address of the recipient but that of the sender. The recipient is free to reply to the email, which lets the two users continue their conversation off-platform.\nForgejo link: design proposal\nI want to hear from you! #Would any of this be helpful to you? Should I work on something else? I am looking to interview FOSS practitioners about those ideas. Book a call with me or get in touch to give me your feedback!\n","date":"19 May 2026","permalink":"https://antonin.delpeuch.eu/posts/improving-organization-management-in-forgejo/","section":"Posts","summary":"\u003cp\u003eIn January, I started a 2-year research project on the governance of FOSS projects.\nI do this as part of the \u003ca href=\"https://aces.telecom-paris.fr/\" target=\"_blank\" rel=\"noreferrer\"\u003eACES team\u003c/a\u003e at \u003ca href=\"https://www.telecom-paris.fr/\" target=\"_blank\" rel=\"noreferrer\"\u003eTélécom Paris\u003c/a\u003e, where I am hosted by \u003ca href=\"https://www.theozimmermann.net/en/\" target=\"_blank\" rel=\"noreferrer\"\u003eThéo Zimmermann\u003c/a\u003e and \u003ca href=\"https://upsilon.cc/~zack/\" target=\"_blank\" rel=\"noreferrer\"\u003eStefano Zacchiroli\u003c/a\u003e. The project is part of the \u003ca href=\"https://www.pepr-ensemble.fr/en/congrats-large-scale-collaboration/\" target=\"_blank\" rel=\"noreferrer\"\u003eCONGRATS area of the eNSEMBLE funding programme\u003c/a\u003e, which looks at the dynamics of online collaboration more broadly.\u003c/p\u003e\n\u003cp\u003eOne part of my project consists in looking into the existing governance dynamics in those projects, trying to identify patterns that work better than others. This involves collecting data from large collections of FOSS projects (via the \u003ca href=\"https://www.softwareheritage.org/\" target=\"_blank\" rel=\"noreferrer\"\u003eSoftware Heritage\u003c/a\u003e archive), interviewing people, and analyzing those results. I\u0026rsquo;ll write more about this later.\u003c/p\u003e","title":"Improving organization management in Forgejo"},{"content":"Lately I\u0026rsquo;ve been thinking about a problem which feels like an interesting governance one. (Yet another, yes!) The recent abandonment of the nvim-treesitter repository made quite some waves, so it feels a fitting occasion to expose this problem here.\nTree-sitter is a parser generator. You specify a grammar for a language (say, Java or CSS) and it generates pretty efficient parsing code for it in C. This parser can then be used in all sorts of applications, such as for syntax highlighting in text editors (Neovim, Helix, Zed, Emacs…). It\u0026rsquo;s also used in GitHub to list function definitions in files, for instance. We also rely on it in Mergiraf, our resolution tool for git conflicts.\nFor each language that you want to parse with tree-sitter, you need a grammar (written in the format understood by tree-sitter), which often comes together with a \u0026ldquo;scanner\u0026rdquo; (custom lexing code written in C). Of course, such parsers require maintenance: for instance because the target language evolves, because bugs are found, to improve performance, or simply because tree-sitter itself evolves. So far, so good: each parser can just be an open-source project of its own, with people maintaining the parser, and each user simply depending on that software, right?\nSadly, not really. Most tools which rely on tree-sitter parsers can\u0026rsquo;t just use tree-sitter parsers as plugins: they need some additional configuration to give some meaning to the syntax trees created by the parsers. For instance, an editor using tree-sitter for syntax highlighting will need to map node types (defined by the parser) to the colors and other styles used in the editor. Platforms like GitHub need to decide which nodes of the trees will be displayed as function definitions in its UI. Tools like mergiraf need additional configuration to determine which nodes it can merge \u0026ldquo;commutatively\u0026rdquo; (meaning that their local order does not matter). This tool-specific \u0026ldquo;glue\u0026rdquo; also needs maintenance, as the parser and the tool evolve. To support those use cases, tree-sitter comes with a notion of \u0026ldquo;queries\u0026rdquo;, which are ways to select nodes of syntax trees. For instance, a text editor can use queries to define which nodes it will display in bold font, to state things in a somewhat simplistic way. A lot of tools use such queries to define this bridge between the parser and their own concepts, but those queries remain tool-specific, so they still require their own maintenance.\nOn top of that, tools rely on different distribution formats for parsers. Some tools require their users to point them to a Git repository containing the source of the parser. Some others require bindings for the parser in a specific language (such as a Rust crate or a Python module), often uploaded to a package registry (Crates.io, PyPI, NPM). Others require a compiled version of the parser as a dynamic library or a WASM file.\nThe nvim-treesitter repository that was abandoned recently is one that stores this glue code between Neovim, a text editor, and hundreds of tree-sitter parsers. Maintaining this glue looks like pretty gruesome work. One of its maintainers has been visibly strained by that work, and it looks like a conflict with a user was the last straw that pushed him to abandon ship (in the form of archiving the repository).\nSo the big question the tree-sitter ecosystem is struggling with right now is how to organize the maintenance of not just the parsers themselves but also this glue that is required for tools using it. People look for alternatives to the \u0026ldquo;monorepo\u0026rdquo; approach of nvim-treesitter, using one repository per language for instance. Beyond choices of repository structures, the question I\u0026rsquo;m interested in is who does this work and how do they organize themselves.\nMotivations to maintain parsers #People get involved in tree-sitter parser/query maintenance for all sorts of reasons:\nfolks who care about a specific parser and for a specific downstream use. For instance, a Neovim user who codes in PHP may want to get involved in the creation or maintenance of the PHP parser, because it translates to improvements to their work environment. They are knowledgeable about the programming language being parsed, perhaps less so about tree-sitter parser development or distribution. folks who care about a specific downstream tool, but no language in particular. For instance, I work on Mergiraf, which relies on a range of tree-sitter parsers. I have made contributions to (and maintain) various tree-sitter parsers, primarily driven by the motivation to make their use in Mergiraf easier. I do that for parsers of languages that I often don\u0026rsquo;t know at all. perhaps more rarely, people who care about a programming language and maintain a tree-sitter parser for it so that their language is well-supported in a range of software development tools. This happens mostly for relatively niche programming languages. They are obviously experts of their language, but might be less familiar with all of the downstream tools relying on tree-sitter parsers (which there is no authoritative list of). perhaps also less frequently, people who are excited by tree-sitter itself as a cool piece of tech and want to make great parsers with it. They might be more knowledgeable about parser optimization and the latest shiny stuff coming from the tree-sitter project itself. This includes tree-sitter maintainers working on flagship parsers, acting as showcases for tree-sitter. In any case, the involvement of those contributors is generally relatively lightweight, driven by the specific motivation they have. Blame them on being selfish if you want, but maintaining a tree-sitter parser is not a particularly creative and fulfilling task. I would say it\u0026rsquo;s perhaps even a rather boring one that doesn\u0026rsquo;t give you a lot of open source \u0026ldquo;street credibility\u0026rdquo; or employability, as far as I can tell.\nFailure of collaboration in parser maintenance #Can we aggregate all this passing interest of people with varying motivations to maintain an ecosystem of parsers and queries?\nIt generally does not work out very well. What often happens is something along the lines of:\nJane creates a parser for FooLang because she wants to use it in Emacs. She spends quite some time doing this - making a parser from scratch is no easy feat! She writes the queries that work for her with her use of Emacs. She publishes her parser as a GitHub repository, under her own account. Three months later, Barnaby opens an issue asking if she could upload her parser to NPM, so that that he can use it more easily in his Node.js project. Jane doesn\u0026rsquo;t have an NPM account, so it\u0026rsquo;s a bit annoying for her to do so (and to commit to publishing the following releases there as well). She declines. Barnaby forks the project and sets up a pipeline to publish it to NPM under his own account. Two months later, Eliott opens a PR to fix a bug on Jane\u0026rsquo;s repository, as he found a case where a valid construct of FooLang isn\u0026rsquo;t parsed correctly. Jane happily accepts, updates the queries accordingly and publishes a new version. Barnaby\u0026rsquo;s fork doesn\u0026rsquo;t get the improvement. Four months later, Jessica opens a PR to Jane\u0026rsquo;s repo, to add support for a new feature of FooLang. Jane has switched to VSCode so she doesn\u0026rsquo;t care so much about this parser anymore and doesn\u0026rsquo;t review the PR. You get the idea: the end state is a network of forks, some of which are distributed in some package registries, some benefiting from various bug fixes in the parser, and, down the line, a lot of duplication of work.\nCommunity Package Maintenance Organizations #Couldn\u0026rsquo;t those people get together in one GitHub organization and maintain parsers and queries collaboratively? There is tree-sitter-grammars, a GitHub organization which is somewhat close to that. But:\nit does not accept new parsers, nor does it seem to accept new members as far as I can tell. I could also not find written governance for it; while there is some effort to have all parsers follow a set of guidelines, there are still big discrepancies between the parsers in terms of the package registries they get published to or the queries they include, for instance. Since I wasn\u0026rsquo;t able to join this organization, I created a similar one called grammar-orchard (on Codeberg), hosting a slowly-growing collection of parsers. There is a lot more I should do to improve this organization, improve the contribution and onboarding process more, advertise the grammars to end users, distribute the parsers in more formats, but it has the merit to exist and to have already enabled nice collaborations. For instance, our Java parser has received sustained contributions from two contributors who reviewed each other\u0026rsquo;s patches after I onboarded them, while the upstream repository doesn\u0026rsquo;t seem to have anyone available to review external PRs. Similarly, I collaborated with a Rust team member on our Rust parser, leading to a lot of improvements to the parser (which in turn is used to \u0026ldquo;fuzz\u0026rdquo; the Rust compiler to avoid crashes). If you are interested in fostering collaboration around your tree-sitter parser, consider joining us and moving your parser there!\nBoth of those organizations can be seen as \u0026ldquo;Community Package Maintenance Organizations\u0026rdquo;, a concept coined by my colleague Théo Zimmermann. He has studied many of those organizations, trying to understand why people start them and what makes them successful. I definitely need to integrate a lot of the lessons learned from his research in the Grammar Orchard!\nThat being said, none of those organizations claim to solve the problem of maintaining this \u0026ldquo;glue\u0026rdquo; between parsers and reusers (which is what the tree-sitter community is currently more concerned with).\nI don\u0026rsquo;t know if such organizations could be extended to also maintain queries or other configuration glue for selected end-user tools. It could be possible to onboard people who have a specific interest in those tools and who help maintain the queries directly in the repositories where the parsers are developed. I don\u0026rsquo;t know to what extent it can work, but it could be worth trying.\nThere is of course also the question to which extent that configuration glue can be mutualized between tools. The fact that those tools have different feature sets, design choices and target audiences gives me little hope that sufficient normalization can happen for this to really bear fruit.\nComments can go to the associated Mastodon post as usual :)\n","date":"16 April 2026","permalink":"https://antonin.delpeuch.eu/posts/the-puzzle-of-tree-sitter-parser-maintenance-and-distribution/","section":"Posts","summary":"\u003cp\u003eLately I\u0026rsquo;ve been thinking about a problem which feels like an interesting governance one. (Yet another, yes!)\nThe recent \u003ca href=\"https://lobste.rs/s/jr4acs/nvim_treesitter_repository_was_archived\" target=\"_blank\" rel=\"noreferrer\"\u003eabandonment of the \u003ccode\u003envim-treesitter\u003c/code\u003e repository\u003c/a\u003e made quite some waves, so it feels a fitting occasion to expose this problem here.\u003c/p\u003e\n\u003cp\u003e\u003ca href=\"https://tree-sitter.github.io/tree-sitter/\" target=\"_blank\" rel=\"noreferrer\"\u003eTree-sitter\u003c/a\u003e is a parser generator.\nYou specify a grammar for a language (say, Java or CSS) and it generates pretty efficient parsing code for it in C.\nThis parser can then be used in all sorts of applications, such as for syntax highlighting in text editors (\u003ca href=\"https://neovim.io/doc/user/treesitter/\" target=\"_blank\" rel=\"noreferrer\"\u003eNeovim\u003c/a\u003e, \u003ca href=\"https://helix-editor.com/\" target=\"_blank\" rel=\"noreferrer\"\u003eHelix\u003c/a\u003e, \u003ca href=\"https://zed.dev/blog/syntax-aware-editing\" target=\"_blank\" rel=\"noreferrer\"\u003eZed\u003c/a\u003e, \u003ca href=\"https://lists.gnu.org/archive/html/emacs-devel/2022-11/msg01443.html\" target=\"_blank\" rel=\"noreferrer\"\u003eEmacs\u003c/a\u003e…). It\u0026rsquo;s also \u003ca href=\"https://github.blog/engineering/architecture-optimization/codegen-semantics-improved-language-support-system/\" target=\"_blank\" rel=\"noreferrer\"\u003eused in GitHub\u003c/a\u003e to list function definitions in files, for instance. We also rely on it in \u003ca href=\"https://mergiraf.org\" target=\"_blank\" rel=\"noreferrer\"\u003eMergiraf\u003c/a\u003e, our resolution tool for git conflicts.\u003c/p\u003e","title":"The puzzle of tree-sitter parser maintenance and distribution"},{"content":"In a recent event where a hackathon was organized, the organizer felt the need to emphasize multiple times that the hackathon wasn\u0026rsquo;t just \u0026ldquo;for crazy programmers who code all night long\u0026rdquo; but was more broadly open to all forms of co-creation between the attendees. It reminded me that I have some problems with this word, which I want to summarize here.\nTo me, the word \u0026ldquo;hackathon\u0026rdquo; suggests sitting down to program something during an in-person event. The goal would be to have a prototype of a system, that would be presented at the end of the event, in some sort of final showcase.\nI find that odd, because it does not feel like a good use of my time in an in-person event. I need quiet, prolongued time and good ergonomics to work on programming tasks. In-person encounters with colleagues are rare enough that I would rather use the time to chat, brainstorm, do exploratory design work for instance. I find that I already have enough of a tendency to start hacky prototypes on a whim, so I don\u0026rsquo;t need the incentive of an event for that.\nThe term can also be associated with competitive vibes (a winning team might be selected at the end of the event). I\u0026rsquo;ve attended many hackathons and none of them were actually competitive, but I imagine that some people might still be put off from the concept because of an expected competitiveness. At least, a \u0026ldquo;hacking marathon\u0026rdquo; smells like extreme performances, energy drinks and sweat, which doesn\u0026rsquo;t feel like the best approach to include less confident folks.\nSo what other words could we use? Just a \u0026ldquo;workshop\u0026rdquo;? A \u0026ldquo;design party\u0026rdquo;? An \u0026ldquo;ideas factory\u0026rdquo;? Maybe some other portmanteau word can be made in that direction?\nThis is squarely my own perspective and I am super grateful to event organizers for putting in their time and energy to organize such events, often on a voluntary basis. The ones I have been to were overwhelmingly well intentioned and very pleasant to attend. Keep them coming! But if you feel like correcting people\u0026rsquo;s expectations every time you use the H word - maybe we can find a better one?\nUpdate: many suggestions of other wordings are in the reactions on Mastodon.\n","date":"4 February 2026","permalink":"https://antonin.delpeuch.eu/posts/is-there-a-better-word-for-hackathon/","section":"Posts","summary":"\u003cp\u003eIn a recent event where a hackathon was organized, the organizer felt the need to emphasize multiple times that the hackathon wasn\u0026rsquo;t just \u0026ldquo;for crazy programmers who code all night long\u0026rdquo; but was more broadly open to all forms of co-creation between the attendees. It reminded me that I have some problems with this word, which I want to summarize here.\u003c/p\u003e\n\u003cp\u003eTo me, the word \u0026ldquo;hackathon\u0026rdquo; suggests sitting down to program something during an in-person event. The goal would be to have a prototype of a system, that would be presented at the end of the event, in some sort of final showcase.\u003c/p\u003e","title":"Is there a better word for 'hackathon'?"},{"content":"What if the main thing your FOSS project needed was that you work less on it? It has definitely been the case for me, and I think I am not the only one.\nWe regularly hear stories of heroic volunteer open source maintainers single-handedly sustaining critical cornerstones of our digital infrastructure. We often hear how those selfless workers have been in that position for decades, tirelessly fixing bugs after work in late night hacking sessions. Many of them are role models in our movement. They define what it is to be successful as a FOSS maintainer.\nWithout disputing the effort and skill that it takes to reach such a position, I want to explore a few reasons why your projects might actually be better off if your time was somewhat more scarce. Of course, situations vary, but I think this applies to a lot of projects, large and small.\nIt feels quite ironic for me to write such advice, since I routinely fail to follow it, but ridicule doesn\u0026rsquo;t kill, so let\u0026rsquo;s dive right in!\nMaking space for other contributors #Here I\u0026rsquo;m assuming that you are interested in having more contributors in your project, which is not necessarily a given.\nThere isn\u0026rsquo;t really a reason for someone to get involved in your project if you are consistently doing what needs to be done there. The speed at which you tackle issues might be too low to your taste, but from the outside of the project it probably still looks like you are managing pretty well.\nWhen considering whether to tackle an issue yourself, it\u0026rsquo;s worth pondering whether it might be worth leaving it to others. If an issue isn\u0026rsquo;t critical, is relatable and is technically approachable for someone not too familiar with the project, then it\u0026rsquo;s likely a good candidate to be used as \u0026ldquo;contributor bait\u0026rdquo;. Make it clear that you support the issue being solved, that you don\u0026rsquo;t have capacity to work on it yourself and that you welcome contributions on it.\nOnboarding contributors is work, of course. Reviewing someone\u0026rsquo;s contribution on that issue might take you more time than solving it yourself. It is rare that contributors stick around, but when they do, the \u0026ldquo;return on investment\u0026rdquo; of your time can be considerable. If you consider this a waste of your time, then it is going to be hard to build a community around your project.\nBy restricting the time you spend on a project that you care about, your own interest in onboarding other contributors might increase, helping you to focus your energy on that, instead of on tackling issues yourself.\nEven if you have already onboarded other contributors, the fact that they have less time available for working on your common project can constitute a significant power imbalance. This is a known issue of do-ocracies, a governance style widespread in FOSS.\nAdjusting priorities #One symptom that you might have too much time on your hands is when you develop the Not Invented Here syndrome. Say you decide to make your own localization system for your project, because clearly all the existing ones don\u0026rsquo;t match your standards. You\u0026rsquo;re not only investing time into that, but you are also increasing the maintenance load and making it harder for others to get acquainted with your project given that it uses non-standard tools.\nEven if you don\u0026rsquo;t reinvent the wheel, being very particular about various aspects of your project that aren\u0026rsquo;t really critical (say, code formatting) is mostly about marking your own territory. Behind the facade of enforcing quality standards, you are primarily asserting your ownership of the project and demonstrating this power to other contributors. Is this nitpicking really a good use of your valuable time?\nMental and physical health #It sounds cliché, but even if you get a kick by working on those things, it\u0026rsquo;s probably healthy to have other hobbies too! I am neither a psychologist nor an occupational health expert, but it should be common sense, no? I feel it\u0026rsquo;s also easier to maintain an enthusiastic and friendly communication when not feeling too stretched. Intuitively, that should have quite some impact on your project too.\nPersonally, I notice that my productivity drops after prolongued work sessions. When I go offline for a week, I\u0026rsquo;m always surprised how little notifications I got when I come back and how well the internet has been able to cope without me. Including my FOSS projects.\nGetting involved in more non-FOSS things is also a great way to inform your FOSS work by getting inspiration from other domains. By the way your local Aikido club is organized. By the chats with a random passer-by during an evening stroll. By the struggles with IT that you observe at the polling station during your local election. It\u0026rsquo;s hard to stay creative if you just evolve in the same online sphere full of likeminded people.\nThat\u0026rsquo;s all I can come up with for now, but maybe you have other ideas? Rotten tomatoes and pitchforks can be directed at this Mastodon post.\n","date":"18 November 2025","permalink":"https://antonin.delpeuch.eu/posts/too-much-time-on-your-hands/","section":"Posts","summary":"\u003cp\u003eWhat if the main thing your FOSS project needed was that you work less on it?\nIt has definitely been the case for me, and I think I am not the only one.\u003c/p\u003e\n\u003cp\u003eWe regularly hear stories of heroic volunteer open source maintainers single-handedly\nsustaining critical cornerstones of our digital infrastructure. We often hear how those selfless workers\nhave been in that position for decades, tirelessly fixing bugs after work in late night hacking sessions.\nMany of them are role models in our movement. They define what it is to be successful as a FOSS maintainer.\u003c/p\u003e","title":"Too much time on your hands"},{"content":"By default, forges like GitHub make you create new repositories under your own account. But most open source projects should rather be in an \u0026ldquo;organization\u0026rdquo;, even if there isn\u0026rsquo;t a corresponding legal entity for them.\nIn this post, we\u0026rsquo;ll assume your repository is on GitHub, but most of it applies to other forges like GitLab or Codeberg.\nThe downsides of personal repos #If you\u0026rsquo;re interested in attracting contributors to your project, personal repositories are not very well suited for that. To onboard trusted contributors, your only option is to add them as \u0026ldquo;Collaborators\u0026rdquo;. This gives them write access to the repository, but not much more than that: they cannot onboard more people themselves, nor can they access most other repository settings.\nBeyond the permissions that you can grant them, various other aspects of the platform will influence their relationship to your project. First, you are the only owner of the repository. Your username is prominently displayed in the header of every page of the repository and in the URL. This means that people also more likely to see you as the (only) person responsible for fixing things when they break, or reviewing external contributions. Second, the \u0026ldquo;Collaborator\u0026rdquo; status that you grant to trusted contributors is almost not advertised at all. The GitHub profile of a collaborator doesn\u0026rsquo;t show that you have granted them this trust, so it\u0026rsquo;s harder for them to take credit for their work on your project. In the GitHub repository itself, the list of collaborators isn\u0026rsquo;t shown either. As far as I know, it\u0026rsquo;s only possible to find out who is a collaborator by looking for messages that they wrote in issues or pull requests, where a small badge will be shown, or by inferring it from other actions (such as them merging a pull request).\nAnother aspect that can make a big difference is that collaborators that you invite do not automatically watch the project. So by default, they do not get notified by new issues or pull requests, making it less likely that they take responsibility for them.\nOnboarding contributors with organizations #Meet GitHub organizations! When a repository is owned by an org, it means that:\npeople can be added to the project with various configurable permission levels, including as co-owners users are able to advertise their membership to the project, shown both on their profile and on the organization\u0026rsquo;s profile depending on their permissions, new members automatically watch certain repositories the project is more identifiable as a team endeavour, as the association to your account is visually less prominent the organization can be used to host multiple repositories that are topically linked. For instance, if you have one repository for a library and one for an example application that uses that library, it makes sense to have them both in the same organization. Migrating is easy! #It takes three simple steps:\ncreate an organization (selecting the free, open source plan). It is common to use the name of the project (which is often the name of the repository) as org name too. go to the settings of your repository and transfer it to the org update any URLs to the repository so that they point to the new address. GitHub sets up redirections anyway, so the change should be transparent for most users. If you had added people as collaborators to the repository, they will remain collaborators of the new repository, but it makes sense to add them as members of the organization instead (so that they can benefit from the above).\nWhy aren\u0026rsquo;t organizations the norm? #I wish forge platforms could create an associated organization for each new repository, by default. Creating a personal repository, for instance for your dotfiles, would still be possible but would need to be chosen explicitly.\nBut GitHub has a strong incentive not to do that. Its revenue model is centered around providing services for companies, so it conveniently uses the org creation page to advertise its paid features for teams:\nIf creating a free organization was done automatically at repository creation, this would likely mean less people sign up for the paid features. This unfortunately reinforces the impression that organizations are for companies or other legal entities, while they are totally suited for projects without an institutional home.\nImplicit feudalism #Curious about the implications of platform design for team dynamics? Check out Nathan Schneider\u0026rsquo;s article on \u0026ldquo;implicit feudalism\u0026rdquo;, the tendency of online services to grant absolute power to the creators of a given space (be it a chat channel, social media group, or an open source project). And if you are really serious about making your project a community endeavour, maybe get a governance model for it?\n","date":"14 October 2025","permalink":"https://antonin.delpeuch.eu/posts/move-your-foss-project-to-an-org/","section":"Posts","summary":"\u003cp\u003eBy default, forges like GitHub make you create new repositories under your own account. But most open source projects should rather be in an \u0026ldquo;organization\u0026rdquo;, even if there isn\u0026rsquo;t a corresponding legal entity for them.\u003c/p\u003e\n\u003cp\u003eIn this post, we\u0026rsquo;ll assume your repository is on GitHub, but most of it applies to other forges like GitLab or Codeberg.\u003c/p\u003e\n\u003ch3 id=\"the-downsides-of-personal-repos\" class=\"relative group\"\u003eThe downsides of personal repos \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#the-downsides-of-personal-repos\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eIf you\u0026rsquo;re interested in attracting contributors to your project, personal repositories are not very well suited for that.\nTo onboard trusted contributors, your only option is to add them as \u0026ldquo;Collaborators\u0026rdquo;. This gives them write access to the repository, but not much more than that: they cannot onboard more people themselves, nor can they access most other repository settings.\u003c/p\u003e","title":"Move your FOSS project to an org!"},{"content":"Years ago, I read The Pull Request Hack, a blog post advocating for a radical way of collaborating on FOSS: \u0026ldquo;Whenever somebody sends you a pull request, give them commit access to your project.\u0026rdquo; The post is really worth a read. More than a decade later, I think it aged well: a lot of FOSS projects would benefit from granting permissions to contributors more proactively.\nYou might argue that the situation has changed: since that blog post, some pretty traumatizing social engineering attacks have happened, such as the XZ Utils backdoor or the event-stream incident. In both of those examples, malicious actors abused write accesses that maintainers granted them willingly, after seeing a track record of well-intended contributions. It can be tempting to react to those attacks by raising the bar for granting privileges in FOSS projects, but I think it\u0026rsquo;s misguided. In both cases, the maintainers were particularly vulnerable to social engineering because they were isolated and overworked.1 Intuitively, growing a project team proactively helps reduce this risk, both by having more eyes to scrutinize things and by those eyes being less strained. In other words:\nIf you don\u0026rsquo;t onboard contributors proactively, then you effectively select project members for their social engineering abilities.\nSo, I want to reflect here on my attempts to practice the Pull Request Hack in the past years. I want to share what I have learned from many experiences, either in the positions of maintainer or of new contributor. Needless to say, those are rules that I unfortunately still often fail to follow myself.\nCommunication is key #Granting project accesses is great, but it\u0026rsquo;s important to explain to the new project member what they are invited to do with it. I have often made the mistake of just sending an invite to the GitHub project (for instance, OpenRefine), without much other form of process.\nFirst, people might just not see the invitation or not understand where it came from. Second, they might not be sure which accesses they\u0026rsquo;ve been granted exactly, and under which conditions they are supposed to use them. Most people are actually pretty careful and won\u0026rsquo;t exercise those rights because they don\u0026rsquo;t want to break things.\nIt can be useful to explain why you are granting those accesses: is it because you generally trust them to use those well if you were to disappear from the project (so, as a means of increasing the bus factor), or is it because you\u0026rsquo;re actively looking for help in a specific area right now? Having clarity on that might increase the chances that people use the permissions you grant.\nGranting meaningful rights #So you granted this contributor full write access to the Git repository for this Python library that you\u0026rsquo;re tired of maintaining. Good job! But if making releases involve manually uploading the library to PyPI and you are the only owner of the package on that platform, then they need you to stay reactive to make those releases. Similarly, if the repository is still stored under your own GitHub account and not under an organization one, they won\u0026rsquo;t be able to invite other contributors themselves.\nIt\u0026rsquo;s completely legitimate (and even advisable) not to immediately grant full owner rights to random contributors. Those initial permissions can still be really heartwarming for newcomers. But more advanced rights are often withheld indefinitely from onboarded contributors, with the understanding that they remain the prerogative of a sort of BDFL. Personally, signs of this dynamic discourage me pretty fast from contributing.\nExpiring memberships #If you\u0026rsquo;re serious about using that Pull Request Hack in your project, then you quickly end up with a lot of project members, most of whom have just contributed a couple of things in the past but moved on to other activities since. The fact that someone did not use the privileges they were granted is not a problem on its own: it\u0026rsquo;s even expected that it makes up the majority of cases. Those cases are just there to enable the occasional miracle of someone taking up the offer and climbing higher up on the contributor ladder.\nBut that has downsides: the list of project members is cluttered and it\u0026rsquo;s hard to know who is actually participating. It\u0026rsquo;s also more risky, as dormant accounts could get compromised and damage the project.\nSo you also need to clean up those members in one way or another. And that\u0026rsquo;s tricky to do well. Ideally, you want to:\nReach out to the contributors before removing them, so that they have a chance of letting you know if they have any interest in contributing again. That\u0026rsquo;s a great occasion to remind them of your project and could lead to retaining them. It also avoids giving them the bad taste of discovering it after the fact from a cold system notification. Have a criterion to decide who to remove, so that it doesn\u0026rsquo;t feel like there is resentment behind it. With bonus points if the criterion is implemented in some system that generates notifications for those expirations. I find it very hard to remove project members on my own initiative, because they are (almost always) valuable contributors that I am thankful to, or even befriended with. So it helps if there is some sort of system supporting me for that. Get a governance model! #All of those principles are a lot easier to follow if you have adopted a governance model for your project. If there is a document describing the different roles people can have in your project and how one can get from one to the other, then:\nCommunicating about the expectations around how permissions are used is a lot easier: you can point people to that document. Granting meaningful rights is a lot easier, because your governance document defines your own role in the project too, and how contributors can expect to reach it (if at all). The criteria for expiring memberships should of course also be documented there, which serves as a useful anchoring, making it clear that cleaning up the project members list is not a personal vendetta against inactive contributors. If you are looking for inspirations, why not take a look at the model we use in Mergiraf? It is tailored to small projects and is easy to adopt. The FOSS Governance Collection contains a lot of useful examples from more established projects. The OSS Watch website also offers useful advice and templates.\nDid I forget other things you should keep in mind when using the Pull Request Hack? Let me know on Mastodon!\nThe interview of the former event-stream maintainer on the Changelog podcast or press coverage of the XZ attack helps understand the social circumstances around the attacks.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","date":"1 August 2025","permalink":"https://antonin.delpeuch.eu/posts/the-pull-request-hack-is-not-enough/","section":"Posts","summary":"\u003cp\u003eYears ago, I read \u003ca href=\"https://felixge.de/2013/03/11/the-pull-request-hack/\" target=\"_blank\" rel=\"noreferrer\"\u003eThe Pull Request Hack\u003c/a\u003e, a blog post advocating for a radical way of collaborating on FOSS: \u0026ldquo;\u003cstrong\u003eWhenever somebody sends you a pull request, give them commit access to your project.\u003c/strong\u003e\u0026rdquo; The\npost is really worth a read. More than a decade later, I think it aged well: a lot of FOSS projects would benefit from granting permissions to contributors more proactively.\u003c/p\u003e\n\u003cp\u003eYou might argue that the situation \u003cem\u003ehas\u003c/em\u003e changed: since that blog post, some pretty traumatizing social engineering attacks have happened, such as the \u003ca href=\"https://en.wikipedia.org/wiki/XZ_Utils_backdoor\" target=\"_blank\" rel=\"noreferrer\"\u003eXZ Utils backdoor\u003c/a\u003e or the \u003ca href=\"https://blog.npmjs.org/post/180565383195/details-about-the-event-stream-incident\" target=\"_blank\" rel=\"noreferrer\"\u003eevent-stream\nincident\u003c/a\u003e. In both of those examples, malicious actors abused write accesses that maintainers granted them willingly, after seeing a track record of well-intended contributions.\nIt can be tempting to react to those attacks by raising the bar for granting privileges in FOSS projects, but I think it\u0026rsquo;s misguided. In both cases, the maintainers were particularly vulnerable to social engineering because they were isolated and overworked.\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e Intuitively, growing a project team proactively helps reduce this risk, both by having more eyes to scrutinize things and by those eyes being less strained. In other words:\u003c/p\u003e","title":"The Pull Request Hack is not enough"},{"content":"Some months ago, I saw a job posting for Wikibase.Cloud\u0026rsquo;s product manager. It piqued my interest and lead me to ask myself what I would do if I were to work at Wikimedia Deutschland, in a position to influence the direction the Wikibase project. I realized I have opinions, so this is an attempt to put them in an understandable format, mostly in the interest of being able to refer to it in discussions with people as those subjects regularly come up.\nExperiment with more federation scenarios #The volume of data stored in Wikidata has been growing steadily for the past few years, and we have known that this growth isn\u0026rsquo;t sustainable. Wikidata\u0026rsquo;s usefulness relies in a big part on its data being made available in its Query Service, a triple store which can\u0026rsquo;t be scaled so easily. The Search Team at the Wikimedia Foundation has recently decided to stop offering all of Wikidata in a single Query Service, splitting the graph in two as the load was not manageable anymore. The SQL database itself has been growing at a concerning rate. And beyond the purely technical aspects, there is also the question of how much data the Wikidata community can curate, as the human effort required to take care of all those entities is considerable.\nAt the same time, users of third-party Wikibase instances struggle, because they basically need to build their knowledge graph from scratch, with tooling that is inferior to what is available on Wikidata. Wikibase has been primarily designed to power Wikidata, a centralized and unique knowledge graph. The idea of multiple Wikibases refering to each other is a very foreign concept that doesn\u0026rsquo;t fit so well with the Wikibase way of doing things (such as refering to items via \u0026ldquo;Qids\u0026rdquo;, not qualified by any sort of domain or prefix).\nIt\u0026rsquo;s urgent both for Wikidata and for third-party Wikibases to make it easier to work across multiple Wikibase instances, so that:\ncertain chunks of Wikidata entities can be split out to separate Wikibase instances, the first obvious candidate being the scholarly articles stored by the Wikicite project, people working in third-party Wikibases can more easily build on top of Wikidata and other Wikibase instances. Whether we want to call those changes \u0026ldquo;federation\u0026rdquo; or not, let\u0026rsquo;s make sure they are addressing those pressing needs. Introducing federation in a platform that has not been designed for it can be very difficult, but I believe some lightweight approaches could already make a big difference. A few years ago, Wikimedia Deutschland has experimented with having a Wikibase reuse the properties of another Wikibase instance, which wasn\u0026rsquo;t judged conclusive. Let\u0026rsquo;s just explore more scenarios!\nTo do so, I would focus on the obvious concrete case at hand: Wikicite. This sub-community of Wikidata is a good candidate for splitting it out to another Wikibase instance, because it curates a relatively well delimited set of entities within Wikidata, and is in itself responsible for the lion\u0026rsquo;s share of the scalability issues. The Wikicite project has also been held back by the growth constraints of Wikidata, forcing it to have a patchy coverage of the subject it models, so I believe it could really thrive if it could free itself from those constraints. My priority would be to sit down with Wikicite participants and try to understand what are the blockers that would prevent them from migrating scholarly articles to a separate Wikibase instance as of today.\nHere\u0026rsquo;s how I expect it would go. The first obvious thing to try (in my opinion) would be to load all scholarly articles in a separate Wikibase instance, keeping other general-purpose items (journals, organizations, humans, topics…) in Wikidata. For this to work, we\u0026rsquo;d need to first import Wikidata properties into the Wikibase instance, strategically choosing their datatypes:\na property like cites work (P2860) would retain its item datatype, since it would be used to link mostly from a scholarly article to another (both present in the same Wikibase instance), a property like author (P50) would get a string datatype, so that it can store a Wikidata Qid and appropriately link to it via formatter URL (P1630) and formatter URI for RDF resource (P1921). The exact details of what type of entity to store in which Wikibase can of course be adjusted: the goal would be to make sure that they are consistent with property domains and ranges, so that the separation can be enforced by the property datatypes. For instance, one could also decide to store entities about people alongside scholarly articles, to enable mass-importing of scholarly profiles from external databases without flooding Wikidata. Local items could be equated with Wikidata items via a dedicated property (as is already customary in many Wikibase instances).\nWhat problems do I expect in such an experiment?\nit\u0026rsquo;s going to be cumbersome to add Wikidata Qids as values of properties like \u0026ldquo;author\u0026rdquo;. They will not be rendered with labels, nor be validated when they are saved. Wikicite is heavily reliant on the Wikidata Query Service and expects all the entities relevant to the project to be present in the same triple store. The Scholia app is a good example of that: it is very difficult to rewrite its queries such that they can still run across two triples stores. We can work on those problems.\nWe make a Wikibase extension which declares a new property datatype. Such properties have Wikidata items as values, with the same auto-complete widget. The labels of the entities refered to are cached in a dedicated table, refreshed whenever those statements change or upon explicit request. The local statements refering to Wikidata items export to RDF as you would expect, with the corresponding Wikidata entity URIs. This provides a transparent way to refer to Wikidata items, imitating the existing UX without needing any architectural changes in Wikibase; We run a Query Service containing both the data from the Wikicite Wikibase instance and relevant parts of Wikidata, fed by two query service updaters. Hopefully, a reduced read/write load would make this more tractable than the Wikidata Query Service. This can be trialled at a small scale. Plugging an off-the-shelf query service updater to Wikidata might also not work so well given the volume of changes, some adaptations might be required there. Surely this is not straightforward, but perhaps still worth trying out?\nDon\u0026rsquo;t invest more in EntitySchemas #How do you find the capacity to work on federation? By stopping to work on EntitySchemas! Wikimedia Deutschland has apparently been spending a lot of effort on EntitySchemas lately, to make them linkable via statements and to give them more of the features that are expected of Wikibase entities. From the outside, it looks like this effort took a rather long-winded path (by not making EntitySchemas real Wikibase entities and re-implementing much of the associated functionality separately), but even if EntitySchemas were turned into proper Wikibase entities, a major issue would still remain: ShEx (which is the language in which EntitySchemas are expressed) is designed to validate RDF data and not data expressed in Wikibase\u0026rsquo;s own data model.\nTo understand why that\u0026rsquo;s a problem, we need to think about the use cases that those EntitySchemas are supposed to eventually enable. The ones I am aware of are:\ngenerating a report showing to what extent a set of Wikibase entities (or a single one) complies with the schema, helping identify quality issues and address them via manual or automated edits, doing the same sort of quality assurance on candidate edits from a batch upload tool such as OpenRefine, so that data modelling issues can be addressed ahead of an import, generating data entry forms (similarly to Cradle or Wikibase Lexeme Forms) that match the data modeling conventions on the instance. This helps users manually input new entities without pre-existing knowledge of the expected properties, and saves time by avoiding the need to input those properties in the first place, documenting data modeling conventions in a standard form, to be read directly in ShEx form by other users. The fact that Wikibase stores its data in its own format, which is very different from RDF, is a significant hurdle for providing a satisfactory user experience around use cases 1, 2 and 3. To validate Wikibase data with ShEx, one needs to translate it to RDF. This translation is by now well established, as it\u0026rsquo;s used to populate the query service, but it\u0026rsquo;s not so easy to explain to end users, and it\u0026rsquo;s a one-way street. By this, I mean that I am not aware of any tool (even third party) to create a new Wikidata item by supplying its RDF representation. To provide an experience comparable to that of the WikibaseQualityConstraints extension, which is able to highlight issues with specific statements directly in the Wikibase UI, one would need to find a way to lift the results of ShEx validation back up to the Wikibase data model, and that\u0026rsquo;s difficult precisely because of this one-way street conversion. One can for sure come up with something that works in basic cases, but in my opinion it\u0026rsquo;s going to be very hard to make it really user-friendly and reliable.\nBeyond the fact that ShEx is designed to validate RDF, another issue is its very broad expressivity, which makes it hard to validate ShEx schemas efficiently at scale. This is also a big hurdle for use cases 1, 2 and 3. Some years ago, a call was held on this topic and the idea of defining a subset of ShEx was floated. This would have the benefit of constraining the expressivity to make implementation tractable and could also be the occasion to introduce additional fields required to generate data input forms (such as labels or placeholders for input fields). This could potentially work, but then if we are to define a new format, why not make it validate data in the Wikibase model directly? And at that point it doesn\u0026rsquo;t really have anything to do with ShEx anymore.\nUse case 4 remains, but looks not so exciting to me. ShEx schemas for Wikibase data are not very readable per se, since Wikibase entity URIs are quite cryptic unless you learn Qid/Pids by heart. If people are enthusiastic about the ShEx syntax despite that, they can easily embed ShEx schemas in wiki pages, potentially via a template to help with linking to an external validator. No custom Wikbase development is required to satisfy this use case.\nFor those reasons, I think investing more effort into EntitySchemas should not be a priority.\nMore dogfooding #One constant struggle in the Wikimedia movement, and probably a lot of other volunteer communities supported by a small paid team, is the gap between the volunteers (Wikidata editors, third-party Wikibase users) and the employees (teams at Wikimedia Deutschland and Wikimedia Foundation). By gap I mean cultural gap, but also gap of priorities because of the drift of viewpoint both parties have on the products. It\u0026rsquo;s something that is difficult to avoid. When hiring for a role, organizations will generally favour hiring someone with proven experience for this particular type of role (possibly in another industry), over someone coming from the grassroots community that they serve. Volunteers may have given a lot of their time to the movement, but are they going to fit in the organizational chart? Are they going to be reliable as employees?\nOf course, there are ways to reduce this gap, by having the two parties communicate better (user research interviews, in-person or online gatherings). Those are used in the Wikimedia movement. Still, in the context of Wikidata and Wikibase, I think there is room for improvement. Over the past years, we have seen a lot of turnover in the project manager positions at Wikimedia Deutschland, with vacant positions being filled by professionals who had little previous involvement in the wiki community as far as I can tell. Getting to know the community and forming a deep understanding of the product takes time, and I think it is a lot more likely to happen if those professionals become direct users of their product.\nSo how about we do more dogfooding? Can we have a Wikibase.Cloud product manager who is actively involved as an administrator of a Wikibase instance on that platform? Not for a toy Wikibase created for testing purposes, but for a Wikibase actively used by a community who has a job to do with the Wikibase. Let them run editathons for that community on their work time. I think it would really help them make informed decisions about the product they are in charge of.\nBeyond that, it would of course be ideal to be able to retain those product managers for longer. That\u0026rsquo;s a difficult goal and there likely isn\u0026rsquo;t one single reason why the previous ones have left their positions. But intuitively, having them develop a closer connection to their user base should help with making their work easier and more fulfilling. As someone who has interacted on a regular basis with teams at Wikimedia Deutschland because of my involvement in OpenRefine and the Wikibase Stakeholder Group, I have found it difficult to maintain a working relationship to those teams, given the rapidly changing faces on the other side of the Zoom call.\nAnd what about reconciliation? #I have been working on OpenRefine for what feels like a long time (and will be leaving the project this year). As part of that, I have been promoting its reconciliation protocol as something that linked open data platforms should implement. That includes Wikibase: I think it would be really useful for Wikibase to implement this protocol directly. So it\u0026rsquo;s perhaps surprising to some readers that I don\u0026rsquo;t rank this as the top priority of Wikibase improvements. I hope this helps appreciate how urgent I think federation improvements are. The house is burning! In my opinion, the Wikidata Query Service should never have been allowed to split: the Wikidata community should have been given much more explicit feedback about what sort of growth can be sustained by the infrastructure, so that they can make more informed decisions about whether certain types of entities should be stored in Wikidata or elsewhere. My understanding is that the Wikidata team sees their role as infrastructure maintainers, who are there to serve the community of volunteers and accommodate with the organic growth that happens there. This is a principle that has worked well for Wikipedia, because the editorial boundaries agreed on by the community make growth much more manageable. Wikidata is a wiki of a very different nature, where users can make legitimate mass imports that can put the infrastructure on its knees. Letting community members find out about the limits of the infrastructure by running into those walls is not doing them a favour.\nAnother reason why I wouldn\u0026rsquo;t rank reconciliation as a top priority is that it\u0026rsquo;s a project that can be tackled by an external team relatively well. We\u0026rsquo;ve been trying to tackle that in the Wikibase Stakeholder Group and have some funding applications pending, so perhaps it might even end up happening, who knows!\n","date":"20 January 2025","permalink":"https://antonin.delpeuch.eu/posts/wikibase-day-dreaming/","section":"Posts","summary":"\u003cp\u003eSome months ago, I saw a job posting for \u003ca href=\"https://www.wikibase.cloud/\" target=\"_blank\" rel=\"noreferrer\"\u003eWikibase.Cloud\u003c/a\u003e\u0026rsquo;s product manager.\nIt piqued my interest and lead me to ask myself what I would do if I were to work at \u003ca href=\"https://www.wikimedia.de/\" target=\"_blank\" rel=\"noreferrer\"\u003eWikimedia Deutschland\u003c/a\u003e,\nin a position to influence the direction the \u003ca href=\"https://wikiba.se/\" target=\"_blank\" rel=\"noreferrer\"\u003eWikibase\u003c/a\u003e project.\nI realized I have opinions, so this is an attempt to put them in an understandable format, mostly in the interest of being able to refer to it in discussions with people as those subjects regularly come up.\u003c/p\u003e","title":"Wikibase day-dreaming"},{"content":"A one-person open source project with a governance model, is that ridiculous? In this blog post I want to convince you it isn\u0026rsquo;t, and it could be key to solving widespread problems in the FOSS (Free and Open Source Software) ecosystem.\nThe problems #The free software movement is successful in many ways, but there are still annoying aspects to it. There are many situations where I feel like there is no shortage of people with goodwill, but we still fail to work together.\nAs a user of open source software, I often need to make a pick from a bunch of roughly equivalent solutions. Why do I need to choose between KeePass, KeePassX and KeePassXC? Why so many Linux distributions? Why all those open source navigation apps? If only their authors could work together to offer one real alternative to Google Maps, as functional and well rounded!\nAs a contributor to FOSS projects, I struggle with getting my changes accepted. My proposals routinely go unreviewed for months, if they ever get reviewed at all. When I consider contributing a change, I need to factor in this risk, so I regularly give up because the backlog of open pull requests left to rot is a red flag. I could imagine helping with maintenance of some of my dependencies, but there is often no way to even apply for that in the first place.\nAs a maintainer, I struggle to attract new contributors to my projects. If I do get some contributions, their authors do not stick around much. I get tired of this maintenance work. I feel lonely and depressed.\nWhy is it that way? There are surely many factors, but to me, a crucial one is cultural. We don\u0026rsquo;t really have the culture of setting up structures for team work. When I start a new open source project, I primarily think about the use case I want to support. I want the tool to be really good: easy to use, reliable, architecturally sound, well tested. I make it open source because I want to make it maximally useful to people. I don\u0026rsquo;t have a plan for its sustainability, because I just assume that if my tool is sufficiently useful to enough people, contributors will somehow come and help out - isn\u0026rsquo;t that the point of open source? Look, I have even added a CODE_OF_CONDUCT.md document to my repository, so I\u0026rsquo;m a nice and polite guy that people should be able to work with!\nOpen source is full of people like me. I have been trained to solve technical problems, and to consider human problems as basically not in my department. To build communities around the projects I am involved in, I essentially hope that we\u0026rsquo;ll just bond over our common love for good software and that the rest will follow. This is the core of the problem I want to tackle: this assumption that team work will just succeed on its own if people come to help. In my experience it\u0026rsquo;s really not the case.\nIt\u0026rsquo;s not a ground-breaking thesis either. The essay \u0026ldquo;The tyranny of structurelessness\u0026rdquo; by Jo Freeman did such a great job (in 1970!) at explaining why formal roles and decision procedures are helpful to make teams function properly, even at a small scale.\nA lot of established open source projects know that, and have indeed formal governance models that shape the day-to-day decisions, onboarding processes and other procedures. The FOSS Governance Collection contains plenty of examples of tried and tested models which were likely instrumental in the popularity of the projects which adopted them.\nThe problem is: none of those governance models seem tailored to small projects that just got published and don\u0026rsquo;t have a community yet. As someone who just started a small FOSS project, I can\u0026rsquo;t start requiring \u0026ldquo;three +1 votes from members of the Project Management Committee\u0026rdquo; to publish a release - I am on my own! So I don\u0026rsquo;t choose a governance model yet and just let the forge platform shape the permissions and interactions in my project.\nThe default governance model that forges push us into #In the absence of a conscious choice of governance model, the default workflows defined by the forge (such as GitHub or GitLab) are the de-facto governance model of the project. Let\u0026rsquo;s see what that looks like.\nWhen I publish a new open source project on GitHub, the repository gets created under my own user account by default. I am the one and only \u0026ldquo;Owner\u0026rdquo; of the repository. I am able to add \u0026ldquo;Collaborators\u0026rdquo; in the settings. They do not have the same privileges as I have: for instance, they are not able to add other collaborators themselves. The list of collaborators on a repository is also not displayed publicly, nor is it possible to apply to become a collaborator via GitHub. If someone wants to help maintain my project, the options they have are fishing for my email address, or communicate by opening an \u0026ldquo;Issue\u0026rdquo;. Isn\u0026rsquo;t that a great metaphor? The fact that they offer their help is treated as an issue by the platform!\nTo be able to add co-owners, one needs to create an organization and tranfer the ownership of the repository to that organization. It does give more governance options, with the ability to create teams and assign rather granular permissions to them, but I am mostly left on my own to define those teams (I need to take the initiative to set them up) and it still does not let people apply to join them.\nTo summarize, in GitHub\u0026rsquo;s default governance model, the project has one leader and they are there to stay. It\u0026rsquo;s called implicit feudalism and corresponds rather well to the notion of Benevolent Dictator For Life (BDFL) used in the free software movement. While some projects make a conscious decision to adopt that model, it isn\u0026rsquo;t very helpful to grow a community and avoid maintainer burnout.\nNeedless to say, the default settings offered by forges have a huge impact on the overall ecosystem. I would be interested in working on improving those defaults, and Forgejo feels like a fitting project where to explore interventions around this problem.\nLicenses and the \u0026ldquo;just fork it\u0026rdquo; mentality #Another reason for the lack of interest in governance models is, I think, the belief that an open source license is the only real governance model a project needs. It goes like this: if you are not happy with the way the project is run, you can \u0026ldquo;just fork it\u0026rdquo; and you have a copy that you can run the way you want. If you just stick to what most licenses say, publishing an open source project isn\u0026rsquo;t a promise to review and integrate other people\u0026rsquo;s changes in it, nor to onboard anyone on the team. Obviously, advising people to fork a project if they\u0026rsquo;re unhappy isn\u0026rsquo;t exactly ideal for community building. There are other issues with this stance, which are by now well known, such as the importance of certain assets (domain names, coordinates in package repositories…) that cannot be retained in a fork, or the difficulty for users to keep an overview of the fork landscape of a project.\nStill, one nice achievement of the open source movement is that we have a common understanding of the importance of licenses. We understand that assigning licenses to software projects is crucial to enable their adoption, both for users and contributors. Many forge platforms will explicitly encourage you to add one when creating a project.\nCan we grow the same sort of awareness for governance models? Can we get to a stage where most open source developers would have the reflex of systematically adding a governance model to their projects when publishing them, even for the tiniest, most insignificant libraries? The first step towards that goal is to convince you, reader, that it would be a worthwhile pursuit.\nWhen I told a friend that I had adopted a governance model for a project I had just published and where I was the only contributor, their reaction was to ask: \u0026ldquo;what is there to govern?\u0026rdquo;. I think that\u0026rsquo;s a pretty natural reaction! So here\u0026rsquo;s my response. I wanted to:\nmake it clear to prospective contributors that they are welcome to get involved, by showing a clear path to co-maintainership, make it clear to myself how I hope to integrate people in the project and make myself eventually redundant, so that I keep this as a goal even at the initial stage of the project where I have a lot of enthusiasm for working on it and taking responsibility for things; set up a basis for collective decision making before the need for it arises, because it\u0026rsquo;s a lot easier that way. The main problem is: the existing FOSS governance models are designed for mature projects and it feels like a lot of effort to design one from scratch. Why should I reinvent the wheel? My project isn\u0026rsquo;t that special, and I am not a governance specialist so I wish I could adopt an off-the-shelf model. Just like I am not a copyright lawyer and wouldn\u0026rsquo;t want to write a new license for every project I publish.\nOff-the-shelf governance models? #So this is what I have been dreaming of lately: when you create a git repository on a forge platform, it nudges you towards adopting a governance model for your project, and proposes a selection of well-known ones that you can add in one click to the repository as a GOVERNANCE.md file. Just like it nudges you to adopt a license. Ideally, it could also set up the appropriate teams and apply other configuration settings implied by the governance model you picked. Unlike licenses, which are mostly meant to be copied verbatim and not modified, the governance model would be meant to be a starter template to be adapted to the needs of the project as it grows (by specifying how the governance model can be changed). An additional benefit of having such well-known starter packs is that it would help prospective users and contributors quickly grasp the general spirit of a governance model, without having to read it all (just like we have built a common understanding of what the MIT or GPL licenses are and we don\u0026rsquo;t need to dissect them every time we interact with a project licensed as such).\nWhat could this initial offering of governance models look like? Here are ideas of a few options:\nthe BDFL model. There is one owner of the project, who does not intend to share ownership with others, but accepts contributions via pull requests, and may give limited privileges to some contributors for them to help out with certain tasks (such as issue triage and pull request reviewing). Although I am not enthusiastic about this model, it still makes sense to offer it as an option, given that it\u0026rsquo;s the current default. As a potential contributor, I would already find it useful to know that a project has consciously adopted this model. the \u0026ldquo;just fork it\u0026rdquo; model. The owner has no interest or capacity for integrating changes from others. The recommended way to improve this software is to fork it. Projects which adopt this model could disable pull requests and/or issues on the repository to make this clearer. This is a model that can make sense in a lot of cases: for instance, when academics publish code alongside an article to make their research reproducible, they rarely have the intention to develop a thriving community of contributors around the repository. Just like a research article is generally meant to be a final artifact that does not evolve after publication, so is often the associated code. Seen as a governance model, this is quite nihilistic, but it would still be worth stating, because it would help potential users and contributors better understand the intent of the authors and avoid wasting time trying to contribute to it directly. what I would call the Kanthaus model. Kanthaus is a collective I have been involved in for a few years. It has a governance model (called \u0026ldquo;Constitution\u0026rdquo;) with a position system defining how people can transition between the three roles (\u0026ldquo;Visitor\u0026rdquo;, \u0026ldquo;Volunteer\u0026rdquo;, \u0026ldquo;Member\u0026rdquo;). This system makes it possible for newcomers to have a clear pathway towards reaching the same rights and responsibilities as people who have originally founded the collective (\u0026ldquo;Member\u0026rdquo; status). The same sort of tiered system can be used in open source projects, and this is what I have been experimenting in Mergiraf, with small tweaks to make the model still applicable to a one-person team. Adopting a model like this makes sense when the creator of the project does not intend to retain full control over it indefinitely, and instead wants it to form a common that others will be able to steer and maintain. a wiki-style model, where write permissions to the code-base are given in a very lax way. Instead of avoiding the introduction of bugs by requiring reviews to code changes, bugs are remediated by reverting problematic changes after they were committed. I have heard of this model but I am not sure where it is used (if at all). It might be the right choice in certain cases. other established governance models, such as the Apache way, which could be used from the start if the project is already supported by a group of actors. Perhaps some can also be adapted to be made relevant for single project authors, still priming the project towards a governance fit for a mature project. Surely there could be many more options! Let me know which ones you\u0026rsquo;d recommend. I am particularly interested in those which are already applicable right at the project start and help grow a contributor community. I have been thinking that surely, such templates must already exist out there, but so far I couldn\u0026rsquo;t find that. What I am aware of is:\nthe FOSS Governance Collection, listing governance documents from mature projects, most of which are tailored to their particular situations and are not designed to be reused in generic FOSS projects, the CommunityRule website, a \u0026ldquo;governance toolkit for great communities\u0026rdquo; (so, targeting a broader audience than the FOSS ecosystem). It does come with \u0026ldquo;templates\u0026rdquo;, but those templates are not formulated in a way that can be applied to FOSS projects directly, articles such as \u0026ldquo;Understanding open source governance models\u0026rdquo; by Red Hat employees, which categorize governance models in broad families. That is helpful to get an overview of the possibilities, but again, it does not directly give me a viable template that I can apply to my project, established governance models like the Apache way or CNCF\u0026rsquo;s governance templates, which are designed to be applied in a specific institutional context: projects which have already reached a certain size and are affiliated to the corresponding foundations. Let me also know if you think this dream is misguided or ill-founded. And if you are interested in working together on such a collection of reusable governance models, I would be thrilled to hear that. Comments can be posted as replies to this Mastodon post.\nThe \"You wouldn't publish a repo with a license\" meme was generated on the awesome youwouldntsteala.website. ","date":"20 December 2024","permalink":"https://antonin.delpeuch.eu/posts/off-the-shelf-governance-models-for-small-foss-projects/","section":"Posts","summary":"\u003cp\u003eA one-person open source project with a governance model, is that ridiculous? In this blog post I want to convince you it isn\u0026rsquo;t, and it could be key to solving widespread problems in the FOSS (Free and Open Source Software) ecosystem.\u003c/p\u003e\n\u003ch2 id=\"the-problems\" class=\"relative group\"\u003eThe problems \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#the-problems\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h2\u003e\u003cp\u003eThe free software movement is successful in many ways, but there are still annoying aspects to it. There are many situations where I feel like there is no shortage of people with goodwill, but we still fail to work together.\u003c/p\u003e","title":"Off-the-shelf governance models for small FOSS projects?"},{"content":"At the end of this year, we\u0026rsquo;ll be discontinuing Dissemin, a web platform designed to help researchers upload their articles to open repositories. I want to take the time to reflect a bit on this journey and what I learned in the process.\nDissemin started at École normale supérieure around 2015, when a bunch of us were motivated to try and do our bit to open up access to scientific publications. The open access movement wasn\u0026rsquo;t a new thing then, even at ENS. Marie Farge, a researcher at the geosciences department of ENS, had been campaigning for open access for quite some time, and Pablo Rauzy, a fellow computer science student a few years before me, had written a very helpful page in French on the topic and gave many talks on the matter. As I read my first scientific articles and also tried to write some, I discovered the problem of lack of access and got in touch with Marie and Pablo. I also discovered the open access movement online, through many different resources such as those of Richard Poynder, Peter Suber and of course the film The Internet\u0026rsquo;s Own Boy about Aaron Swartz\u0026rsquo;s work, which I found very impressive and inspiring.\nOne thing that the open access movement had been advocating for was that universities and other research institutions should adopt \u0026ldquo;open access mandates\u0026rdquo;, which ask the researchers in those institutions to make their articles openly accessible. The idea was that beyond exercising some form of authority on researchers, such mandates would actually help them get better terms from the publishers they submit their articles to (because publishers need to comply if they want to stay in business). So we somehow set out to adopt such a policy at ENS. Patricia Mirabile and I got elected as student representatives on the scientific council of ENS and after some back and forth managed to get the ENS to adopt such a policy. Hurray! Well, except that nobody would actually care (or even know of) that policy, so we didn\u0026rsquo;t expect a lot of change from that. Normally, institutions which adopt such mandates have some sort of team from the university library which takes care of spreading the gospel among researchers and helping them leverage the policy for their own articles. Because the mandate came from a student initiative, we didn\u0026rsquo;t have that.\nSo, that\u0026rsquo;s where the initial interest in a web platform came: what if we had a website which could list the publications of researchers at ENS and give an overview of the state of their openness? And perhaps even ease the process of uploading them to open repositories? I started working on such a platform and was quickly joined by amazing teammates who were really kind not to be put off by my pretty dire lack of experience and the effect it had on the initial code base. It was pretty disastrous. I will spare you the software engineering lessons learned, because there are way too many and they aren\u0026rsquo;t particularly original. I just assumed I knew a lot of things which I really didn\u0026rsquo;t. We also shifted focus away from providing this service to just one institution (ENS) to making a platform that would work for any researcher, which was in a sense much more ambitious, but it also sidestepped the tricky problem of delimiting what ENS\u0026rsquo; research output was. And we founded a non-profit association to host the platform and continue its development.\nDespite our generous efforts, it never really got anywhere in terms of usage. Some people did upload some articles through Dissemin into open archives, but the numbers weren\u0026rsquo;t exactly impressive. To me, there were two main problems, which lead me to stopping work on the project in 2017.\nThe first problem was the premise of the open access movement, that all academic papers should be made accessible. To me, it felt self-evident: it is in the interest of society and of the researchers themselves to have the result of their work distributed as broadly as possible, so that we can all learn from their findings. My very brief attempt at being a researcher myself convinced me that unfortunately, articles are rarely written to be read. They are instead written to be reviewed and accepted for publication (which should indeed imply a few people reading it in the process, yes, but that\u0026rsquo;s a pretty small readership). In other words, the main motivation for researchers is often to get the recognition and validation of their work, not so much to tell the world about what they did. Communication can happen through papers, but also though a lot of other means: talks, blog posts, courses, popular science magazines, podcasts… whereas papers are the primary token for career progress. Many PhD programmes have explicit requirements about the number of papers (and perhaps the rankings of the journals they appear in) needed to obtain a degree, for instance. And that\u0026rsquo;s fair: of course, it\u0026rsquo;s a dumb productivity metric, but all productivity metrics are. My conclusion is that it\u0026rsquo;s perhaps not worth campaigning to make all papers accessible, if they are not meant to be read in the first place. If a researcher cares about being read, they have many options.\nThe second problem, specific to Dissemin, is that it wasn\u0026rsquo;t designed to cater for any actual user need. The motivation to build this platform was our outrage at the lack of access, and the urge to do something about it. My mental model for who would use Dissemin was very blurry: a researcher, encouraged by their institution, would go on the platform to check which of their articles weren\u0026rsquo;t accessible, and upload those to open repositories. But that\u0026rsquo;s a really bizarre use case: it would be a platform they would be constrained to use by their hierarchy. They wouldn\u0026rsquo;t spontaneously need the platform or see any immediate benefit. So it\u0026rsquo;s no surprise it wasn\u0026rsquo;t used much.\nOn top of that, the landscape evolved since we first started working on this project. The French national open repository, HAL, improved its deposit process so that people don\u0026rsquo;t have to input so much metadata on their own. One of the selling points of Dissemin was that researchers wouldn\u0026rsquo;t have to enter again the metadata of their articles to upload them, as that metadata was already fetched from our data sources. But having this done by the repository itself has obvious advantages over running an additional platform like Dissemin, which adds confusion for users. Also, gold open access became more of a norm, reducing the need for self archiving.\nIn any case, I really don\u0026rsquo;t regret working on this project. I had a great time and made great friends. And I learned that throwing an app at a social problem isn\u0026rsquo;t necessarily the best approach.\n","date":"24 November 2024","permalink":"https://antonin.delpeuch.eu/posts/retrospective-of-dissemin-and-thoughts-on-the-open-access-movement/","section":"Posts","summary":"\u003cp\u003eAt the end of this year, we\u0026rsquo;ll be discontinuing \u003ca href=\"https://dissem.in\" target=\"_blank\" rel=\"noreferrer\"\u003eDissemin\u003c/a\u003e, a web platform designed to help researchers upload their articles to open repositories. I want to take the time to reflect a bit on this journey and what I learned\nin the process.\u003c/p\u003e\n\u003cp\u003eDissemin started at \u003ca href=\"https://www.ens.psl.eu/\" target=\"_blank\" rel=\"noreferrer\"\u003eÉcole normale supérieure\u003c/a\u003e around 2015, when a bunch of us were motivated to try and do our bit to open up access to scientific publications. The open access movement wasn\u0026rsquo;t a new thing then, even at ENS. \u003ca href=\"https://en.wikipedia.org/wiki/Marie_Farge\" target=\"_blank\" rel=\"noreferrer\"\u003eMarie\nFarge\u003c/a\u003e, a researcher at the geosciences department of ENS, had been campaigning for open access for quite some time, and \u003ca href=\"https://pablo.rauzy.name/\" target=\"_blank\" rel=\"noreferrer\"\u003ePablo Rauzy\u003c/a\u003e, a fellow computer science student a few years\nbefore me, had written \u003ca href=\"https://pablo.rauzy.name/openaccess.html\" target=\"_blank\" rel=\"noreferrer\"\u003ea very helpful page in French on the topic\u003c/a\u003e and gave many talks on the matter. As I read my first scientific articles and also tried to write some, I discovered the problem of lack of access\nand got in touch with Marie and Pablo. I also discovered the open access movement online, through many different resources such as those of \u003ca href=\"https://poynder.blogspot.com/\" target=\"_blank\" rel=\"noreferrer\"\u003eRichard Poynder\u003c/a\u003e, Peter Suber and of course the film \u003ca href=\"https://en.wikipedia.org/wiki/The_Internet%27s_Own_Boy\" target=\"_blank\" rel=\"noreferrer\"\u003eThe Internet\u0026rsquo;s Own Boy\u003c/a\u003e about \u003ca href=\"https://en.wikipedia.org/wiki/Aaron_Swartz\" target=\"_blank\" rel=\"noreferrer\"\u003eAaron Swartz\u003c/a\u003e\u0026rsquo;s work, which\nI found very impressive and inspiring.\u003c/p\u003e","title":"Retrospective of Dissemin and thoughts on the open access movement"},{"content":"A central feature of Git is the ability to merge the contents of diverging revisions. It underpins not just the git merge command, but also rebase, cherry-pick and revert for instance. Without it, no collaboration would be possible. And it generally works great.\nOne limitation is that it\u0026rsquo;s line-based. If the two sides touch neighbouring lines, we need to manually resolve a merge conflict. That happens even if the changes are touching independent syntactic elements, because Git\u0026rsquo;s merging heuristic doesn\u0026rsquo;t know about syntax. Which begs the question: what if it did?\nThere has been research on this topic and various prototypes of syntax-aware merging algorithms exist out there. Git offers an extension point to provide a custom \u0026ldquo;merge driver\u0026rdquo;, an executable taking the two diverging versions of a file together with their common ancestor and producing the merged file. Curiously, this extensibility does not seem to be used that much. I couldn\u0026rsquo;t find any open source merge driver that would do syntax-aware merging for common programming languages, and be reliable enough for daily use. I tried improving Spork, the most promising one I could find, but I quickly ran into issues that couldn\u0026rsquo;t be fixed easily in the existing code base.\nSo I made one: it\u0026rsquo;s called Mergiraf. It was fun to make and I hope it can be useful to others.\nI have tried to put the focus on usability, in contrast to academic prototypes which generally focus on quantitative evaluations on benchmarks. That means:\nbeing fast enough for interactive use: I don\u0026rsquo;t want to slow down my daily git commands for the sake of solving some conflicts once in a while, erring on the side of caution by producing conflicts markers in tricky cases, as an incorrectly solved conflict can be difficult to detect and have serious consequences, offering ways out of situations where Mergiraf produces an incorrect merge (with utilities to review its work and report issues easily), because those can and will happen, comprehensive documentation written as a user manual, not a research paper, dog-fooding: I have been using it for a while for real development on other projects and it\u0026rsquo;s working well so far. So, would you use this? I suspect most people don\u0026rsquo;t encounter enough conflicts in their daily work to bother installing it. But I think Mergiraf can be quite useful if you maintain a fork and regularly synchronize it with upstream, in which case conflicts often pop up. I actually find it helpful even to reorganize my own work, making it quite a bit easier to reorder commits on a branch for instance. Avoiding conflicts with myself, in a sense.\nI have also tried setting up Mergiraf as a project that\u0026rsquo;s inviting to new contributors. For instance by writing a detailed tutorial about adding support for a new language and a GOVERNANCE.md file setting out how people can get involved (Is it presomptuous to have a governance model for a tiny project where I\u0026rsquo;m the only contributor? More on that later.)\nMany thanks to the Spork team for their help understanding and tinkering Spork, to various friends who provided various sorts of feedback at various stages, and to Freya F-T for the illustrations in the documentation!\nFeedback welcome - for instance as issues on Codeberg.\n","date":"9 November 2024","permalink":"https://antonin.delpeuch.eu/posts/mergiraf-a-syntax-aware-merge-driver-for-git/","section":"Posts","summary":"\u003cp\u003eA central feature of Git is the ability to merge the contents of diverging revisions. It underpins not just the \u003ccode\u003egit merge\u003c/code\u003e command, but also \u003ccode\u003erebase\u003c/code\u003e, \u003ccode\u003echerry-pick\u003c/code\u003e and \u003ccode\u003erevert\u003c/code\u003e for instance. Without it, no collaboration would be possible. And it generally works great.\u003c/p\u003e\n\u003cp\u003eOne limitation is that it\u0026rsquo;s line-based. If the two sides touch neighbouring lines, we need to manually resolve a merge conflict. That happens even if the changes are touching independent syntactic elements, because Git\u0026rsquo;s merging heuristic doesn\u0026rsquo;t know about syntax. Which begs the question: what if it did?\u003c/p\u003e","title":"Mergiraf: a syntax-aware merge driver for Git"},{"content":"Welcome to my eighth contribution experience report. See the one about Git for some background about the initiative.\nSpork is a tool to merge diverging versions of Java source files. It\u0026rsquo;s the result of a Masters\u0026rsquo; thesis and has a corresponding scientific article published about it, so it clearly identifies as research software.\nIt\u0026rsquo;s debatable whether such type of open source project is meant to include external contributions at all. A published research paper is normally treated as a final research output, which isn\u0026rsquo;t meant to evolve (apart from possible corrections to fix serious issues), so arguably any accompanying software should be equally final. For reproducibility purposes at least, the paper should point to a precise version of the software with which experiments can be reproduced. But there are also many examples of major open source projects which started off as research artifacts and got a life of their own after that.\nMy motivation to contribute #I had written my own little tool to fix import conflicts when merging Java files, which worked okay, but I was interested to see if I could instead use an existing tool, more principled and powerful. Spork seemed to be the state of the art, so it felt like a good candidate. Trying it out, I found various issues, and making small patches to address them felt like a refreshing distraction from working on OpenRefine.\nFirst contact with the project #Opening an issue about the lack of compatibility with Java 21, I got helpful explanations from the original authors about the difficulties in offering such a support. Other issues about improvements to the algorithm itself were similarly well received - it felt like a very fitting channel of communication to have in-depth discussions and learn from the authors\u0026rsquo; experience.\nDevelopment environment #It\u0026rsquo;s a Java project so I was used to the tooling. The fact that it mixes Kotlin and Java in the same project was a bit of a hurdle, I had to switch to using IntelliJ instead of Eclipse because out-of-the-box support for this set-up was better there.\nFinding my way into the code base #It\u0026rsquo;s a relatively small code base so that wasn\u0026rsquo;t much of an issue.\nTesting infrastructure #Testing is mostly done via an end-to-end suite, with many examples of different merge scenarios. It feels fitting for a tool like this.\nReviewing experience #Just like the discussion on issues, the review feedback was very helpful. For instance we had a pleasant joint investigation of the shortcomings of GraalVM which produces native executables which don\u0026rsquo;t crash but behave differently to running the tool on a standard Java virtual machine. But on another pull request, it turned out that the main developer had lost the mental context necessary to review my improvements to the algorithm - which is definitely understandable given the academic aspect of the project.\nCode formatting #I don\u0026rsquo;t remember any particular style being enforced.\nGovernance and roadmap #It feels odd to talk about governance and roadmap for a research prototype, right? Still, I think it would generally be useful to know if authors have the intention of developing an open source project beyond the publication of the associated article, or if they have moved on. That\u0026rsquo;s one form of governance, answering the basic question: is this project open to contributions at all? Or is it better to improve on this project by making a separate code base, with an associated research paper? Not sure if it ought to be formalized, though.\nAfter the main developer signalled he couldn\u0026rsquo;t review one of my contributions, another team member reached out to ask if I\u0026rsquo;d be up for joining the team and continuing development of the project, which I definitely appreciated.\nWould I contribute again? #Despite the good experience on the human level, no, because I don\u0026rsquo;t think the prototype can be turned into a properly usable tool for various reasons. The algorithm is tightly coupled to the Spoon Java parser, which restricts it to merging Java files only, despite the fact that it shouldn\u0026rsquo;t be very hard to generalize the abstract algorithm to other languages. The reliance on reflection inside this parser also makes it hard to turn it into a native binary, which feels like a must to me if it is to be used as a Git merge driver.\n","date":"8 November 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-spork/","section":"Posts","summary":"\u003cp\u003eWelcome to my eighth contribution experience report. See the one about \u003ca href=\"/posts/contribution-experience-report-git\"\u003eGit\u003c/a\u003e for some background about the initiative.\u003c/p\u003e\n\u003cp\u003e\u003ca href=\"https://github.com/ASSERT-KTH/spork\" target=\"_blank\" rel=\"noreferrer\"\u003eSpork\u003c/a\u003e is a tool to merge diverging versions of Java source files. It\u0026rsquo;s the result of a Masters\u0026rsquo; thesis and has \u003ca href=\"https://arxiv.org/abs/2202.05329\" target=\"_blank\" rel=\"noreferrer\"\u003ea corresponding scientific article\u003c/a\u003e published about it, so it clearly identifies as research software.\u003c/p\u003e\n\u003cp\u003eIt\u0026rsquo;s debatable whether such type of open source project is meant to include external contributions at all. A published research paper is normally treated as a final research output, which isn\u0026rsquo;t meant to evolve (apart from possible corrections to fix serious issues), so arguably any accompanying software should be equally final. For reproducibility purposes at least, the paper should point to a precise version of the software with which experiments can be reproduced. But there are also many examples of major open source projects which started off as research artifacts and got a life of their own after that.\u003c/p\u003e","title":"Contribution experience report: Spork"},{"content":"Welcome to my seventh contribution experience report. I have done others for:\nGit, where I give some background about the initiative Nextcloud\u0026rsquo;s docker image Forgejo JGit Mattermost Organic Maps My motivation to contribute #Sometimes I forget to join meetings. Generally without being able to offer any good excuse for it. But three years ago I skipped one and I had a genuine reason: it was Thunderbird\u0026rsquo;s fault! It failed to import an invite in my calendar. When I tried to import it again, it gave me a cryptic error: \u0026ldquo;Processing message failed. Status: 80004005.\u0026rdquo; After some digging I found out that it was because the UID field of the ICS file contained some special characters. Frustrating!\nFirst contact with the project #Back then, I clearly had the urge to do something about it, since I took the time to write a bug report about it and even try to contribute a patch. But I somehow couldn\u0026rsquo;t get Thunderbird to compile on my machine. It was all really heavy and complicated, so I gave up. The patch was left to rot as an attachment to the Bugzilla ticket.\nThree years later, someone commented on the ticket that they had the same problem, and this comment landed in my inbox (in Thunderbird, of course). I was pleased to see that I was not alone. Someone else indirectly confirmed that my sloppiness with meetings isn\u0026rsquo;t entirely my own fault, what a relief! I wanted to try contributing again.\nDevelopment environment #So I searched for \u0026ldquo;thunderbird contribute\u0026rdquo; online and found this beautiful landing page. I think it does the job very well, offering many ways to contribute without being overwhelming either. From there it was easy to find my way to the Linux build prerequisites.\nThis felt so much better than last time I tried! I mean, you still need a lot of resources to build Thunderbird, but they have spent time working on a bootstrap.py script which takes care of setting up just about everything for you. And it tells you what it does and asks you for permission when necessary, as far as I can remember.\nwget https://hg.mozilla.org/comm-central/raw-file/tip/python/rocboot/bin/bootstrap.py chmod +x bootstrap.py ./bootstrap.py I did run into an error at some point, so it wasn\u0026rsquo;t completely flawless:\n*** failed to set up extension firefoxtree: b\u0026#39;wrapfunction target name should be `str`, not `bytes`\u0026#39; Traceback (most recent call last): File \u0026#34;/usr/lib/python3/dist-packages/mercurial/extensions.py\u0026#34;, line 268, in _runextsetup extsetup(ui) File \u0026#34;/home/antonin/.mozbuild/version-control-tools/hgext/firefoxtree/__init__.py\u0026#34;, line 664, in extsetup extensions.wrapfunction(hg, b\u0026#39;_peerorrepo\u0026#39;, peerorrepo) File \u0026#34;/usr/lib/python3/dist-packages/mercurial/extensions.py\u0026#34;, line 678, in wrapfunction raise TypeError(msg) TypeError: b\u0026#39;wrapfunction target name should be `str`, not `bytes`\u0026#39; But I could fix this by removing a .mozbuild directory likely left over from my previous attempt 3 years ago. After a compilation time of about 30 minutes, I had a freshly baked Thunderbird (I say baked because of the heat generated by my processor over the course of the build time).\nAnyway, congrats to the Thunderbird team for improving those set-up instructions, it made a real difference to me.\nFinding my way into the code base #Since I had crafted a patch three years earlier, I luckily didn\u0026rsquo;t need to figure it all out from scratch. The changes were in a .jsm file, which contains what looks a lot like Javascript. I applied my patch and thankfully the new compilation round only took a few seconds, so I could easily confirm that my patch indeed solved the issue. So let\u0026rsquo;s try and get this accepted upstream!\nTesting infrastructure #Expecting that a test would be required to merge such a change, I tried to figure out how the code I changed was currently tested, by grepping for various class or method names from the class I modified. This led me to a related test class, where I could duplicate an existing test case and adapt it to my needs.\nThe tests came under various flavors (cached and uncached) and I didn\u0026rsquo;t really know what it meant. I went for making an uncached one. It felt reassuring to check that the test case failed without my fix and passed with it.\nFor this process I was very lucky to discover that I could run single test files with:\n./mach xpcshell-test comm/calendar/test/unit/providers/test_caldavCalendar_uncached.js There are so many projects where it\u0026rsquo;s actually painful or even borderline impossible to run a single test case from the command line (hello multi-module Maven projects!), it really bugs me. It\u0026rsquo;s such a basic need! I am grateful I didn\u0026rsquo;t have to set up a full-blown IDE with support for JSM files (whatever that looks like) to get the privilege of running a single test case. Good job again, Thunderbird team!\nReviewing experience #To get this reviewed, I somehow needed to submit this to Phabricator, for which I first needed to setup 2FA. Fine.\nThe process of submitting the patch via Mercurial didn\u0026rsquo;t feel familiar at all, but it was well documented. The hardest part was to figure out who to pick as a reviewer, as they have to be mentioned in the commit message. The documentation even states it: \u0026ldquo;It can be pretty tricky to figure out who to ask for a review\u0026rdquo;.\nI tried going to https://wiki.mozilla.org/Modules/Calendar. Back then it had a list of people, but only with full names and not usernames, which were used in commit messages. I had to dig through the commit log to try to manually match those names to commit identities and then reviewer ids (which are different things!). After a lot of guesswork I settled on requesting a review from \u0026ldquo;darktrojan\u0026rdquo;. That process felt really unnecessarily complicated. But by now, there is a new page listing code owners with full names and usernames together! Amazing!\nTo do the actual submission, I tried:\n$ moz-phab submit Failed to find draft commits to submit The failure was due to the fact that I was in the Firefox repository (yes, Thunderbird is built by cloning the Firefox repository and changing some config files to transmutate a fox into a bird). Doing it from the comm/ subfolder worked.\nI enjoyed the fact that I could do hg commit --amend to fix my mistakes, and that running moz-phab submit again detected the changes and updated the patch accordingly. It felt like forgiving tooling.\nThe time to first review was just 2 hours, with friendly feedback requesting sensible changes. Once my patch \u0026ldquo;landed\u0026rdquo;, as they say, they quickly took the initiative to backport it on a release branch, so my changes got released pretty quickly (and I think it was even an LTS release). Very lucky me!\nCode formatting #This is perhaps the best part. I didn\u0026rsquo;t format my changes properly, but they did it for me! Why isn\u0026rsquo;t that the case more often? It\u0026rsquo;s such a waste of time to dedicate a round of review to ask the contributor to run whatever formatting tool on their side and resubmit. Nicely done, Thunderbird!\nGovernance and roadmap #I could find a decription of various trust levels for code contributors at Mozilla, which seems to apply to Thunderbird. It feels a bit scary but it does describe some sort of a pathway to become a trusted with merge rights, at least on a formal level. That\u0026rsquo;s a fairly basic requirement but many projects don\u0026rsquo;t even meet that standard, so I guess that\u0026rsquo;s good. Concerning the roadmap, I am vaguely familiar with their plans by following Thunderbird on Mastodon.\nWould I contribute again? #Yes! Despite the scary codebase, the experience was a good one. And the impact of contributing to such a project feels undeniable to me. We need good email clients if we want to preserve what\u0026rsquo;s left of the federated nature of email, to resist against the Gmail / Office365 oligopoly.\n","date":"17 October 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-thunderbird/","section":"Posts","summary":"\u003cp\u003eWelcome to my seventh contribution experience report. I have done others for:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-git\"\u003eGit\u003c/a\u003e, where I give some background about the initiative\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-nextclouds-docker-image\"\u003eNextcloud\u0026rsquo;s docker image\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-forgejo/\"\u003eForgejo\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-jgit/\"\u003eJGit\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-mattermost/\"\u003eMattermost\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-organic-maps/\"\u003eOrganic Maps\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"my-motivation-to-contribute\" class=\"relative group\"\u003eMy motivation to contribute \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#my-motivation-to-contribute\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eSometimes I forget to join meetings. Generally without being able to offer any good excuse for it. But three years ago I skipped one and I had a genuine reason: it was Thunderbird\u0026rsquo;s fault!\nIt failed to import an invite in my calendar.\nWhen I tried to import it again, it gave me a cryptic error: \u0026ldquo;Processing message failed. Status: 80004005.\u0026rdquo;\nAfter some digging I found out that it was because the \u003ccode\u003eUID\u003c/code\u003e field of the ICS file contained some special characters. Frustrating!\u003c/p\u003e","title":"Contribution experience report: Thunderbird"},{"content":"Welcome to my sixth contribution experience report. I have done others for:\nGit, where I give some background about the initiative Nextcloud\u0026rsquo;s docker image Forgejo JGit Mattermost My motivation to contribute #In Saxony, many bars allow indoor smoking. As someone who struggles in smoky places, I try to avoid those bars. It\u0026rsquo;s often not so easy to look up the smoking status of a place, as it is rarely advertised online, so for a while I have been surveying bars and adding this information on OpenStreetMap, so it can be found easily. Well, easily for me, but not that easily for most people around me, because OpenStreetMap.org isn\u0026rsquo;t user-friendly. So it\u0026rsquo;s not easy to share the result of my evening walks with them.\nI generally don\u0026rsquo;t use a smartphone to go to places, but that\u0026rsquo;s the norm for most people. I had the impression that the Organic Maps app is the most user-friendly way to use OpenStreetMap for the general public. Because this app does not yet display the smoking status of places, I thought I would try to add that.\nFirst contact with the project #I first wrote an issue on GitHub to propose the feature, making a mock-up of what it could look like. The feedback was supportive and I was encouraged to make a further mock-up of what the editing experience could look like. So it felt like it was worth giving it a shot!\nDevelopment environment #I had zero experience with mobile app development (and in fact, very little with mobile app use) so this was a pretty interesting dive. The development environment is reasonably well documented and rightly warns you about the heavy resources needed to start developing. Not only is the repository itself large but so are the pieces of tooling (such as Android Studio, a tweaked version of IntelliJ IDEA for Android development), dependencies, and maps data. This heavy set up felt like a real hurdle, but I guess it\u0026rsquo;s probably forced on them by the mobile app ecosystem. The Android and iOS app are developed in the same repository, as different codebases which share some tooling and localization. They don\u0026rsquo;t have exactly the same features but seem to be quite similar. The repository also contains many other tools, primarily to generate the map files consumed by the apps, which are compiled out of OpenStreetMap extracts and other data sources. All those separate code bases have some intrinsic coupling: for instance, the binary format of the compiled maps needs to be synchronized between the apps and the generators, so that\u0026rsquo;s probably one reason to have them all in the same repository.\nFinding my way into the code base #The first step to implement my feature was to include the smoking=* OSM tag in the compiled maps, by modifying the maps generator written in C++. Not knowing the code base at all I thought I would just imitate other features of points of interest, such as the cuisine=* tag indicating which sort of food is served in a place, or the wheelchair=* tag which classifies wheelchair accessibility. It involved quite a bit of trial and error, working exclusively using the test suite of the generator to validate my changes as I wasn\u0026rsquo;t sure how to inspect the binary output itself. It felt pretty much like a shot in the dark but I was counting on the code review to validate the approach. In particular, I was initially very confused by the difference between \u0026ldquo;types\u0026rdquo; and \u0026ldquo;metadata\u0026rdquo;. It seems that the former makes the information searchable, whereas the latter is just bits of information added to a node without any index, but I wish I could have read that in some documentation (if that exists?).\nOne oddity I noticed is the development workflow for localization: beyond the source files where translations are stored, there are other files that are derived from them and which are checked into the repository. A Python script is included to update those generated files, with the understanding that they should not be edited manually. Intuitively, those files should rather not be checked in at all and generated on the fly before building. That would get rid of a lot of commits (look for \u0026ldquo;[strings] Regenerated\u0026rdquo;).\nReviewing experience #I first went for adding smoking information as a \u0026ldquo;type\u0026rdquo;, not \u0026ldquo;metadata\u0026rdquo;, since it could be useful to search for a places based on their smoking status (like Osmand allows). The first review feedback was that I should rather add this as metadata, which turned out to be much simpler. The reviews came quickly (multiple reviews on the day the PR was opened) and felt supportive.\nBut then it became unclear whether this smoking status had its place in the app at all, with concerns being voiced about the usefulness of the information globally, or that its presence in the app would encourage people to smoke. I spent some time doing a survey of the usefulness of this information country-by-country in an attempt to make the case for including this information at least in some places. But even with that, it seemed that there was no clear consensus for including this status. The verdict came: you can work on it, but we might pull it out. So I gave up.\nOne contributor called for a holistic re-design of the UI which shows information about places, reviewing which information should be included there and under what form. I agree that it would be very useful, given that this panel is currently quite rough on the edges. It omits various types of useful information and the information it displays is often not very clear, or needlessly takes a lot of space. Adding new metadata fields in a piecemeal fashion is unlikely to improve that. But it\u0026rsquo;s unclear to me which process should be followed for such a re-design, and it feels like expanding the scope of my contribution quite a bit.\nTesting infrastructure #The test suite for the C++ was good enough for my purposes as I could just imitate the surrounding context to write tests for my changes.\nFor the Android app, I could not find any tests, which I found quite curious. I don\u0026rsquo;t know if there are end-to-end testing frameworks for Android, but surely some parts could be covered by unit tests at least?\nCode formatting #For the generator in C++, I found CPP_STYLE.md which documents the expected style with quite some detail. According to it, I can use clang-format to format my code according to the guidelines. But running clang-format -i indexer/feature_data.cpp seems to butcher everything, which isn\u0026rsquo;t great.\nFor Java code, I initially thought it would be handled automatically by the IDE, but sadly no, the way it is enforced is by relying on reviewers to manually add comments in the PR about the code style, which doesn\u0026rsquo;t feel like a great use of everyone\u0026rsquo;s time.\nGovernance and roadmap #I could not find any document about the governance of the project. I have been told it is in docs/GOUVERNANCE.md.\nI did not find a page about a roadmap either, but the general direction of the project seems relatively clear to me: the state objective seems to be a viable FOSS competitor to Google Maps, with a focus on privacy and cleanliness from undesirable app features.\nWould I contribute again? #Probably not, given the various problems I encountered.\n","date":"13 July 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-organic-maps/","section":"Posts","summary":"\u003cp\u003eWelcome to my sixth contribution experience report. I have done others for:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-git\"\u003eGit\u003c/a\u003e, where I give some background about the initiative\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-nextclouds-docker-image\"\u003eNextcloud\u0026rsquo;s docker image\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-forgejo/\"\u003eForgejo\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-jgit/\"\u003eJGit\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-mattermost/\"\u003eMattermost\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"my-motivation-to-contribute\" class=\"relative group\"\u003eMy motivation to contribute \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#my-motivation-to-contribute\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eIn Saxony, many bars allow indoor smoking. As someone who struggles in smoky places, I try to avoid those bars. It\u0026rsquo;s often not so easy to look up the smoking status of a place, as it is rarely advertised online,\nso for a while I have been surveying bars and adding this information on OpenStreetMap, so it can be found easily. Well, easily for me, but not that easily for most people around me, because \u003ca href=\"https://openstreetmap.org/\" target=\"_blank\" rel=\"noreferrer\"\u003eOpenStreetMap.org\u003c/a\u003e\nisn\u0026rsquo;t user-friendly. So it\u0026rsquo;s not easy to share the result of my evening walks with them.\u003c/p\u003e","title":"Contribution experience report: Organic Maps"},{"content":"Welcome to my fifth contribution experience report. I have done others for:\nGit, where I give some background about the initiative Nextcloud\u0026rsquo;s docker image Forgejo JGit This one is about contributing to Mattermost. Mattermost is a chat application for teams, often seen as a FOSS replacement for Slack as it offers a very similar user experience.\nMy motivation to contribute #I identified a small issue in the user interface: a checkbox not being properly associated with the text that describes its function, meaning that clicking on the text would not have any effect on the checkbox itself.\nI encountered this problem when setting up a test instance of Mattermost to evaluate the feasibility of using it as a Slack replacement for Kanthaus. The problem was clearly not a blocker for me but it felt like a good opportunity to evaluate the contribution process to this project.\nFirst contact with the project #I just went ahead and opened an issue on the mattermost/mattermost GitHub repository. As the Mattermost website makes it clear, this project is run with the \u0026ldquo;open core\u0026rdquo; model, meaning that it is controlled by a company which sells a proprietary version built on top of the open source core. So it wasn\u0026rsquo;t surprising to me that contributing required signing a Contributor License Agreement (enabling them to change the license of the open source project as they please). That\u0026rsquo;s of course not ideal - it\u0026rsquo;d be much better to have the assurance that the core project indeed remains open source and be run by an independent organization. But for the sake of evaluating the experience, I went ahead and signed it.\nDevelopment environment #The change I wanted to make to fix the problem was really simple: adding a containing \u0026lt;label\u0026gt; tag to properly link the checkbox to the text. It\u0026rsquo;s something I could test out by opening the web developer tools in my browser and manually editing the HTML structure. So for the first version of my PR I didn\u0026rsquo;t even set up any sort of development environment and just manually edited the React component where this checkbox was, so that it renders the correct DOM structure.\nWhen more changes were requested I did follow the contributor guide to set up my development environment. That relatively simple and easy to follow. Most of the time was spent waitaing for NPM and Go packages to download. The most cumbersome part of the workflow was that their setup uses docker compose (not docker-compose), and installing it required adding a Debian repository of Docker.com to my sources.list. I always find it nicer when the development tools required aren\u0026rsquo;t so bleeding edge and are readily available in Debian.\nFinding my way into the code base #The technique of searching for UI text (in this case the checkbox\u0026rsquo;s label) worked very well: it directly took me to the React component where the form was defined.\nReviewing experience #My PR got a first review from a developer the next working day, which was nice.\nThe checkbox in question had another problem: it was implemented as a button element, whose appearance would change after being clicked by inserting an SVG image in it to make it look checked. This bizarre set-up was probably introduced to customize the appearance of the checkbox, but it was really weird because (at least on my platform) the appearance was really close to the native one.\nThe reviewer noticed that and asked me to clean this up in the same go. It felt reasonable since it would contribute to cleaning up the DOM structure there and making the form more accessible, so I did this small change (which required me to actually set up a development environment, but it was a bit bold to skip that part anyway).\nAnother interesting aspect was that the developer who first responded added tags to my pull request, marking it as requiring a \u0026ldquo;dev review\u0026rdquo;, a \u0026ldquo;UX review\u0026rdquo; and a \u0026ldquo;QA review\u0026rdquo; - all done by different people!\nThe UX review person noticed a small change in alignment between the checkbox and its label and asked me to bring it back in line with the original. I was given a very precise zoomed-in screenshot with added guide lines, which felt pretty precise and helpful.\nWhat I found very interesting is that they have a system in place to spin up a test Mattermost server from the PR branch, just by adding a label to the pull request. This is something I have considered doing for OpenRefine for some time, because I think it would massively help with getting more community members review pull requests from a user standpoint, without having to learn anything about Git or installing any sort of developer environment on their side. It encourages me to look more into it.\nAll in all it took 19 days from the opening of the PR to it being merged (with me prodding them only once to get another round of reviews).\nUpdate: A few weeks after my PR, I recieved a Mattermost mug in the post, with a little text thanking me for my contribution (printed on the mug). I also got an invite to the \u0026ldquo;Mattermost Developer Meeting\u0026rdquo;. Fascinating experience! Of course I didn\u0026rsquo;t actually need any of those, but it\u0026rsquo;s still gratifying. I can imagine it helps with retaining contributors.\nTesting infrastructure #For my change, no test was requested from the reviewers, so I didn\u0026rsquo;t look into that topic at all.\nCode formatting #The CI did notice some style violations, which I fixed by running make fix-style. On my machine, this command took about 120 seconds to run, which felt not ideal since that\u0026rsquo;s something you generally run often.\nGovernance and roadmap #Mattermost being open core, I just expect the project to be run by the company. So I wouldn\u0026rsquo;t be able to have much influence on the big picture as a contributor. But to fix some small consensual problems, it would probably work.\nThey do have a roadmap which shows the main features they are working on, although in the past rather than in the future as of today. They seem to do a minor release every month.\nWould I contribute again? #If I were to contribute something again to Mattermost core it would probably be something similarly small. I did notice that they have a very nice plugin system (which I find inspiring for OpenRefine too) so if I do more Mattermost development in the future, it would more likely be in that area (for instance if the migration from Slack to Mattermost reveals missing integrations).\n","date":"2 May 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-mattermost/","section":"Posts","summary":"\u003cp\u003eWelcome to my fifth contribution experience report. I have done others for:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-git\"\u003eGit\u003c/a\u003e, where I give some background about the initiative\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-nextclouds-docker-image\"\u003eNextcloud\u0026rsquo;s docker image\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-forgejo/\"\u003eForgejo\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-jgit/\"\u003eJGit\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThis one is about contributing to \u003ca href=\"https://mattermost.com/\" target=\"_blank\" rel=\"noreferrer\"\u003eMattermost\u003c/a\u003e. Mattermost is a chat application for teams, often seen as a FOSS replacement for Slack as it offers a very similar user experience.\u003c/p\u003e\n\u003ch3 id=\"my-motivation-to-contribute\" class=\"relative group\"\u003eMy motivation to contribute \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#my-motivation-to-contribute\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eI identified a small issue in the user interface: a checkbox not being properly associated with the text that describes its function, meaning that clicking on the text would not have any effect on the checkbox itself.\u003c/p\u003e","title":"Contribution experience report: Mattermost"},{"content":"Welcome to my fourth contribution experience report. I have done others for:\nGit, where I give some background about the initiative Nextcloud\u0026rsquo;s docker image Forgejo This episode is the first one about contributing to a library. JGit is a Java implementation of Git, covering many of the features of the original implementation in C. This makes it possible to use Git in Java programs without having to use separate processes and avoids license compatibility issues (Git itself is under GPLv2, JGit uses the Eclipse Distribution License which is similar to the new BSD license). JGit is used in pretty established projects such as the Eclipse IDE or Gerrit.\nMy motivation to contribute #In short: I wanted to fix a bug in the JGit library that I encountered while tweaking a git merge driver.\nIn my contribution report about Git, I explained my interest in custom merge drivers to support my work on OpenRefine. Writing my own little merge driver was easy enough to get rid of merge conflicts in import statements, but I figured out it would likely be more efficient to use an existing merge driver for Java files, to avoid re-inventing the wheel. Spork is the one that looks the most mature as far as I could tell, but it lacks support for the \u0026ldquo;diff3\u0026rdquo; conflict presentation mode, which I rely on heavily to understand how to solve conflicts. So I tried adding support for this in Spork. In some cases Spork falls back on a textual merge implemented in JGit. While writing tests for Spork I encountered a case where JGit itself generated an incorrect merge output. This is the bug I tried to solve in JGit.\nFirst contact with the project #This was a rather messy process. To report the bug, I first landed on Eclipse\u0026rsquo;s Bugzilla instance and was super confused about being asked to \u0026ldquo;select a classification\u0026rdquo; for my bug. I had no idea which classification the JGit project belonged to. I looked at existing JGit bugs in Bugzilla for inspiration, but could not find a classification there either.\nIt somehow occurred to me that all those bugs were pretty old, with barely no recent activity. Via JGit\u0026rsquo;s support page I indeed found a link to the GitHub issues for JGit, which seem to be the new official bug tracker. So I opened issue #38 about my bug.\nI also realized I needed to sign the Eclipse Contributor Agreement to submit patches, which I did. I don\u0026rsquo;t remember the exact experience but it felt rather baroque. My notes say: \u0026ldquo;OMG DCO\u0026rdquo;.\nDevelopment environment #As usual, the first step to contributing to a new project is generally to clone its git repository, and that\u0026rsquo;s straightforward, right? Well, not in this case.\nI first landed on \u0026ldquo;https://git.eclipse.org/r/plugins/gitiles/jgit/jgit\u0026rdquo;, which back then said \u0026ldquo;repository moved to https://git.eclipse.org/r/jgit/jgit\u0026rdquo;. Which was odd because \u0026ldquo;https://git.eclipse.org/r/jgit/jgit\u0026rdquo; returned a bare \u0026ldquo;Not found\u0026rdquo; error in plain text. Still, things seemed to be actively reviewed in Gerrit, so I thought I\u0026rsquo;d try to use that. I found the documentation for Gerrit in Eclipse and followed it to set up my SSH key for upload.\nPushing my changes gave:\n! [remote rejected] HEAD -\u0026gt; refs/for/master (prohibited by Gerrit: project state does not permit write) Not great! I then checked the CONTRIBUTING.md file and found a link to a new Gerrit instance, with more recent activity: https://eclipse.gerrithub.io/q/project:eclipse-jgit/jgit+status:open. That sounds promising! But in the same file, they also still pointed to https://bugs.eclipse.org/bugs/enter_bug.cgi?product=JGit to open issues, which says \u0026ldquo;Sorry, entering a bug into the product JGit has been disabled.\u0026rdquo; Okay, then let\u0026rsquo;s make a patch for that first…\nI went on to switch to the new Gerrit instance. Do I need to add my SSH key again? No, it\u0026rsquo;s handled already when creating my account on GerritHub, fine.\nI did yet another change of git remote with git remote set-url origin \u0026quot;ssh://wetneb@eclipse.gerrithub.io:29418/eclipse-jgit/jgit\u0026quot; Ah, I also forgot about the DCO so I need to do git commit --amend -s.\nFinally, I ran the magical git push origin HEAD:refs/for/master and success, I got a link back: https://review.gerrithub.io/c/eclipse-jgit/jgit/+/1177977\nThat was a really convoluted process, which took me a lot of time to figure out. I think I landed in the project at a rather infortunate time, when they had just made the switch to a new forge and the documentation wasn\u0026rsquo;t quite up to date. Maybe forges need better ways to indicate that the project has moved to a different space.\nIt doesn\u0026rsquo;t help that because the bug tracker is now on GitHub, it would intuitively make sense to be able to submit pull requests there. Alas, pull requests apparently cannot be disabled on GitHub projects, so you need to resort to pull request templates to indicate where to submit changes. Despite that, there is currently an open pull request with what looks like a rather good contribution, which the author abandonned because they could not figure out how to submit it. Not ideal!\nTo summarize, I considered in total four different places where to submit my changes:\nhttps://git.eclipse.org/r/jgit/jgit, now 404 https://git.eclipse.org/r/plugins/gitiles/jgit/jgit which used to point to the above https://github.com/eclipse-jgit/jgit where PRs cannot be disabled https://eclipse.gerrithub.io/q/project:eclipse-jgit/jgit, the right one Of course it makes sense for JGit to use Gerrit, which I have nothing against. There are some pretty nice features that I miss in GitHub, such as showing the merge conflicts with other pending patches.\nWith all this, I still haven\u0026rsquo;t made code changes! Let\u0026rsquo;s open the IDE. This is a project of the Eclipse Foundation, used by the Eclipse IDE itself, so you\u0026rsquo;d think that working on it via Eclipse should work rather well, no? Think again! I imported it in Eclipse successfully, but then noticed that Eclipse makes changes to its project settings which are tracked in the git repo. This means dozens of tracked files with unstaged changes, which can easily slip into a commit if using git commit -a. Maybe I\u0026rsquo;m not using the right version of Eclipse, or my environment isn\u0026rsquo;t configured correctly in another way? In any case that\u0026rsquo;s not very pleasant to work with.\nFinding my way into the code base #Given that I noticed the bug from another project using this library, I already knew the entry point through which the bug was manifesting itself, so from there it was relatively easy to trace it back to the problematic code. The algorithm to merge text files is rather fiddly but I found JGit\u0026rsquo;s implementation to be still fairly readable, especially in comparison to Git itself.\nReviewing experience #My Gerrit patch didn\u0026rsquo;t attract a lot of attention initially. I reached out by email to the person who implemented diff3 support in JGit to have their opinions on my change, but did not get a reply. After two weeks, I tried to rope in some random folks I saw were reviewing other patches. I could eventually attract a reviewer who was kind enough to approve the patch even if they were not familiar with this part of the code base.\nActively pinging people to find reviewers for my contributions is something I often do in such contexts. I feel a bit bad about asking for more work from maintainers who are generally already quite stretched. I try to do it in a kind way, without pressuring people or making them feel bad about the delays. I guess the etiquette around this varies from project to project. The XZ incident did shine some light on the social dynamics at play in such situations.\nTesting infrastructure #I am used to working with Eclipse, so it was pretty simple to run unit tests interactively from the IDE. The particular algorithm I wanted to fix was covered by a pretty good test suite, which I could imitate to add my own tests.\nHowever, running the entire test suite from the command line with mvn clean test fails on my machine:\n[INFO] Results: [INFO] [ERROR] Failures: [ERROR] RacyGitTests.testRacyGitDetection:57 expected:\u0026lt;...de:100644, time:t0, [length:1, content:a][b, mode:100644, time:t0, length:1], content:b]\u0026gt; but was:\u0026lt;...de:100644, time:t0, [smudged, length:0, content:a][b, mode:100644, time:t0, smudged, length:0], content:b]\u0026gt; [INFO] [ERROR] Tests run: 5676, Failures: 1, Errors: 0, Skipped: 110 The test suite is known to have flaky tests like this one, which is a bit sad. Not being familiar with those tests and the code they cover, I\u0026rsquo;d just propose to mark them as \u0026ldquo;expected failures\u0026rdquo;, so that they don\u0026rsquo;t stand in the way of new contributors, but I suspect this wouldn\u0026rsquo;t be accepted as it would be better to fix the tests themselves. I don\u0026rsquo;t have the capacity to start investigating race conditions in those tests, as this can be quite time-consuming.\nCode formatting #Eclipse automatically reformats code when I save, which is nice. I assume this must conform to the project\u0026rsquo;s settings. The downside of this integration is the problem I mentioned earlier of Eclipse updating its configuration files tracked in Git.\nGovernance and roadmap #Looking at the mailing list, the project is actively looking into bringing more people and requiring two reviews before merging contributions. Maybe they\u0026rsquo;ll be interested in this blog post.\nConcerning gouvernance, I am not completely sure how it works. The Eclipse Foundation has general guidelines about the gouvernance of the projects they host, but JGit\u0026rsquo;s own Gouvernance page only lists the releases, which seems rather unrelated. The Getting Involved page does not say much about that either.\nWould I contribute again? #JGit a library that\u0026rsquo;s quite far upstream from my own projects, so it\u0026rsquo;s difficult to justify the time to contribute to it, but if I end up having more free time, I guess it could be fun to make some more contributions. In a sense, the fact that it\u0026rsquo;s a re-implementation of an existing tool means that you don\u0026rsquo;t need to make too complex design decisions, which can be a nice thing. The rather tedious contribution experience I encountered so far is a bit chilling, but maybe it would get better now that I passed those initial hurdles.\n","date":"20 April 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-jgit/","section":"Posts","summary":"\u003cp\u003eWelcome to my fourth contribution experience report. I have done others for:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-git\"\u003eGit\u003c/a\u003e, where I give some background about the initiative\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-nextclouds-docker-image\"\u003eNextcloud\u0026rsquo;s docker image\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-forgejo/\"\u003eForgejo\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eThis episode is the first one about contributing to a library. \u003ca href=\"https://www.eclipse.org/jgit/\" target=\"_blank\" rel=\"noreferrer\"\u003eJGit\u003c/a\u003e is a Java implementation of Git, covering many of the features\nof the original implementation in C. This makes it possible to use Git in Java programs without having to use separate processes\nand avoids license compatibility issues (Git itself is under GPLv2, JGit uses the Eclipse Distribution License which is similar to the new BSD license). JGit is used in pretty established projects such as the \u003ca href=\"https://eclipseide.org/\" target=\"_blank\" rel=\"noreferrer\"\u003eEclipse IDE\u003c/a\u003e or \u003ca href=\"https://www.gerritcodereview.com/\" target=\"_blank\" rel=\"noreferrer\"\u003eGerrit\u003c/a\u003e.\u003c/p\u003e","title":"Contribution experience report: JGit"},{"content":"Welcome to my third contribution experience report. I have done others for:\nGit, where I give some background about the initiative Nextcloud\u0026rsquo;s docker image My motivation to contribute #In short: I wanted to fix some bugs I had encountered, and I was curious about contributing to a \u0026ldquo;soft fork\u0026rdquo;.\nHere is more background. In Kanthaus, we used to rely on GitLab to host about 20 Git repositories for various documents (governance, minutes of meetings), for website sources (such as for https://kanthaus.online or https://handbook.kanthaus.online) or some pieces of software used in the house (such as scripts to steer our heat pump or to gather statistics about the house).\nAlthough we like publishing a lot of those things in the open, some repositories need to be private, for instance if they contain personal information or minutes of meetings about sensitive topics. A GitLab pricing change meant that we could no longer have private repositories for free there, unless we reduced the number of members of our GitLab organization to maximum 5. For a while we kept using it, by setting up a communal account whose credentials were shared in the community. This was obviously not ideal from a security standpoint and it also meant that we couldn\u0026rsquo;t reliably keep track of who did what.\nSelf-hosting GitLab looked difficult given that it seemed to require a lot of resources. Once Forgejo / Gitea developed a continuous integration feature, that started to look like a credible replacement. So I started experimenting with it and encountered a few problems. I wanted to avoid requiring people to create accounts for this service by reusing Nextcloud for authentication, but in some cases this generated internal errors in Forgejo. Also, although Forgejo had support for migrating repositories from a wide range of sources (including GitLab), migrating some of our repositories there failed with a SQL error.\nFirst contact with the project #I first opened an issue about Forgejo\u0026rsquo;s website, to encourage adding a more prominent link to the documentation. The suggestion was well received and although I wrote that I was open to submitting a PR for it, it quickly got implemented by someone else. It felt like a good start!\nDevelopment environment #I had never written Go code before so I did not know how to set up a development environment for this project. I did not intend to do a lot of work on Forgejo so I just tried editing the Go source files with my text editor. I discovered that the Docker image could be built via Docker itself, without having to install any particular tooling on the host machine, thanks to the \u0026ldquo;buildx\u0026rdquo; Docker plugin. Although that takes much longer than \u0026ldquo;normal\u0026rdquo; compilation, it felt attractive not to have to worry about the build dependencies at all. To submit my changes to the project I did need to install some tooling in the end, to run the tests and format my code. If I remember correctly, installing the golang Debian package was sufficient for that.\nFinding my way into the code base #For the first bug I wanted to solve (about OAuth integration), I could find the place where the error was thrown by enabling the development mode on our Forgejo instance. From there, imitating the surrounding code was good enough to implement the fix I wanted. For the second bug it was a little harder. Although the bug was known (in Gitea\u0026rsquo;s bug tracker), the root cause was not understood (at least it was not clear from the issue). I think I investigated it by logging all SQL requests made during the migration of my GitLab repository, but I am not sure if I could do that from Forgejo itself or if it was on PostgreSQL\u0026rsquo;s side. Once I had unterstood the problem it was relatively easy to find where to make my fixes.\nReviewing experience #This is where Forgejo being a \u0026ldquo;soft fork\u0026rdquo; made the experience quite interesting. What they mean by \u0026ldquo;soft fork\u0026rdquo; is that the commits that diverge away from upstream (Gitea) are regularly rebased on top of upstream. This has a lot of consequences for contributors like me.\nFirst, because all Gitea commits eventually make it into Forgejo, I had a choice of submitting my changes to either projects. The reason why I had started using Forgejo and not Gitea is that it looked like Forgejo had put more thought into governance. The folks behind it started the fork because they disagreed about the lack of separation between Gitea as an open source project and as a commercial service provider. As I could relate to that concern, it felt natural to use Forgejo. I also appreciated the fact that Forgejo was developed using Forgejo itself, whereas Gitea\u0026rsquo;s repository is hosted on GitHub, which felt a bit ironic.\nIn this context, submitting my pull request to Forgejo felt more natural: otherwise, I\u0026rsquo;d just be interacting with the Gitea project and Forgejo\u0026rsquo;s governance improvements would be of little use to me. Another important factor was that Forgejo\u0026rsquo;s pull request backlog was much smaller than Gitea\u0026rsquo;s, so it felt like my contributions were more likely to be reviewed swiftly in Forgejo.\nThe review process in Forgejo was really great: I got blazingly quick feedback, people were very supportive and helped improve my not very idiomatic Go. I guess there is a particularly strong incentive to be nice to newcomers when you launch a fork, to gather the critical mass needed to make the fork viable.\nHowever, I have later realized that even if I only care about Forgejo, there is still a strong case for submitting contributions to Gitea. Indeed, by submitting pull requests to Forgejo, I am adding to the stack of commits that need to be rebased each week on top of Gitea, which is a significant burden for the project. If Gitea makes changes to the lines of codes my patches touch, then my patches will need to be rewritten as part of this rebase process. This is a bit of a weird thing to do: in a sense, the person doing the rebase will be putting words in my mouth. This case did actually happen when Gitea wrote a different fix for my second issue (about migration from GitLab). The person doing the rebase did ask for my review when that happened, which was classy, but it does feel like a really complicated process, especially because the rebase must happen quite quickly and atomically, so the project can\u0026rsquo;t really afford lengthy reviewing rounds in such a process.\nAnother consequence of this weekly rebase is that if you have pull requests to Forgejo that are open for longer than a week, you\u0026rsquo;ll need to rebase them yourself on top of the newly-rebased Forgejo. It\u0026rsquo;s not a big deal, but still, it would be nicer without. On the plus side, this encourages everyone to merge pull requests swiftly, I guess.\nNote that the Forgejo project is currently considering to become a hard fork instead.\nTesting infrastructure #Forgejo has a test suite, mostly consisting of unit or integration tests focusing on the backend, written in Go. For my first fix, I was lucky to find that the area I was working on already had a test which I could duplicate and adapt pretty easily to cover my changes.\nFor the second, it was a bit trickier. Because migrating a GitLab repository to Forgejo involves making a lot of HTTP requests to GitLab, those requests need to be mocked to make the test pass reliably. There were already some tests for the GitLab integration which did that, but the use case I needed to test involved making quite a few requests, so it did not feel like it would be really doable to set up all the required mock requests by hand. There was also a test which would run against the live GitLab.com instance, but only if an API key was provided in the test environment. Neither Gitea nor Forgejo provided such an API key in their CI, meaning that the test had not been run for a long time: it was actually failing, because it hadn\u0026rsquo;t been adapted for some recent changes in the code. Not great. I went on to update the test so that it passes, then turn it into a properly mocked test by capturing all the HTTP trafic it generated and turning that into mocked HTTP responses, and finally use the same machinery to write a test for my fix. So there was quite a bit of clean up work needed there.\nFor this testing work it felt very useful to be able to run a single test and not the entire test suite every time. To me this is really a crucial feature in any project I work on, because the machines are work on are typically not very fast and will take many minutes to run the test suite of a project like OpenRefine or Forgejo. I asked on the project chat how people did that and I was surprised to find that it was not that easy (it required crafting a command to invoke the test runner manually), so I took this opportunity to investigate and document how to do that in VSCodium instead.\nCode formatting #When I submitted my first pull request the CI complained about some format violations. I could fix them with the gofmt command, which was the one raising the error in the CI. I then realized I can use make lint to format everything. Also, switching to VSCodium helped to make sure the formatter is run in the background when I save a file.\nGovernance and roadmap #As mentioned before the governance model of Forgejo is pretty thorough (which is to be expected given that this was the motivation to fork). They have a dedicated repository to document their rules, the composition of teams, the decisions they make, and so on. They were also quite proactive in pulling me in, first by adding me to the Contributors team without me even requesting it.\nBecause that did not let me merge pull requests as such, that prompted them to create a new team in their governance model (\u0026ldquo;Mergers\u0026rdquo;) and encourage me to apply to it (even though my Go is clearly quite wobbly, and I hadn\u0026rsquo;t requested the right to merge at all). That\u0026rsquo;s a pretty extraordinary thing to see in a FOSS project.\nIn OpenRefine I have tried to be similarly proactive in pulling people in but I have realized that because our governance model is pretty messy, it would be worth cleaning that up first.\nAlso, as a new team member in Forgejo I had the same sort of weird experience as I expect new OpenRefine contributors have. I do not remember getting any notification from the forge about me being added to the team (in GitHub, you would at least get an invitation by email), and then I got tons of notifications from various repositories in the organization because I started watching them automatically. This is an experience I would be interested in improving: maybe by making more changes to Forgejo, for instance to notify people when they get added to teams (ideally, explaining why they are added, what privileges it grants them, how they are invited to use them, and so on).\nWould I contribute again? #For sure, it was a great experience. It gives me the confidence that I should be able to fix other problems in Forgejo as I encounter them in our use of the platform. I am excited to see how the project evolves, especially with the prospect of a hard fork.\n","date":"21 January 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-forgejo/","section":"Posts","summary":"\u003cp\u003eWelcome to my third contribution experience report. I have done others for:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-git\"\u003eGit\u003c/a\u003e, where I give some background about the initiative\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/posts/contribution-experience-report-nextclouds-docker-image\"\u003eNextcloud\u0026rsquo;s docker image\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"my-motivation-to-contribute\" class=\"relative group\"\u003eMy motivation to contribute \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#my-motivation-to-contribute\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eIn short: I wanted to fix some bugs I had encountered, and I was curious about contributing to a \u0026ldquo;soft fork\u0026rdquo;.\u003c/p\u003e\n\u003cp\u003eHere is more background. In \u003ca href=\"https://kanthaus.online\" target=\"_blank\" rel=\"noreferrer\"\u003eKanthaus\u003c/a\u003e, we used to rely on \u003ca href=\"https://about.gitlab.com/\" target=\"_blank\" rel=\"noreferrer\"\u003eGitLab\u003c/a\u003e to host about 20 Git repositories for various documents (governance, minutes of meetings), for website sources (such as for \u003ca href=\"https://kanthaus.online\" target=\"_blank\" rel=\"noreferrer\"\u003ehttps://kanthaus.online\u003c/a\u003e or\n\u003ca href=\"https://handbook.kanthaus.online\" target=\"_blank\" rel=\"noreferrer\"\u003ehttps://handbook.kanthaus.online\u003c/a\u003e) or some pieces of software used in the house (such as scripts to steer our heat pump or to gather statistics about the house).\u003c/p\u003e","title":"Contribution experience report: Forgejo"},{"content":"Welcome to my second contribution experience report. It follows the one for Git which gives some background about the initiative.\nMy motivation to contribute #In short: I had been hit by a very annoying bug and I wanted to fix it.\nHere is more background. In Kanthaus we have a NextCloud instance that we use for shared documents. I have recently started helping out to maintain it, which means doing regular upgrades.\nAs the instance was a couple of versions behind, I tried updating it to the latest version by changing the tag of the Docker image we use to the latest available version. The Docker image detected that an upgrade from a previous version was necessary, so it started the upgrade process. However, NextCloud only supports upgrading from one major version to the next: jumping major versions is not supported. As a result, the upgrade process failed. What was particularly frustrating however, was that when restoring the previous version, NextCloud refused to start because the data files were apparently already tagged with the more recent version. So, I was left in a state where the NextCloud instance could not start anymore, no matter which version of NextCloud I used. And the backups I had were not fresh enough to afford restoring them. I should obviously have made backups just before the upgrade and checked the manual to realize that the upgrade path was not supported, but that did not make it less frustrating.\nFirst contact with the project #I went to the GitHub repository for the Docker image I used (nextcloud/docker), found the files that define the Docker image and started tweaking them to add an earlier check, to prevent any upgrade from starting if the versions differ too much.\nDevelopment environment #Given that the Docker entry point is simply defined by a bash script, I could edit that with my text editor directly. I don\u0026rsquo;t really enjoy writing Bash, it feels super brittle and easy to break things unintentially, but well, for such a small change, I can hold my nose (and I guess Bash is probably a de-facto standard for such entrypoint scripts anyway).\nFinding my way into the code base #I did not have to look into NextCloud\u0026rsquo;s own code base, just the Docker deployment, so it was a fairly simple job of looking for where the error message I had seen in my migration log was output in the bash script. From there, I could just backtrack to a place where to add my additional condition.\nReviewing experience #It was a bad experience. When opening my pull request, I first tried to collect all the issues it would fix - because the problem had been reported many times: at least in #1809, #1129, #616 and #617.\nWhile making this inventory, I found another pull request filed nearly three years before, doing exactly the same fix (of course, with a different way to compare the versions, but that\u0026rsquo;s a detail). This pull request had been left without any comment and rejected two years later without explanation. That felt particularly frustrating. I had been badly hit by a really annoying problem that was not only well known but also for which a fix had been submitted, and simply left to linger for no apparent reason.\nI pinged someone who seemed to have push access to the repository according to the recent activity on the repository and after two weeks, my PR got reviewed and merged.\nTesting infrastructure #There were some GitHub Actions running on my PR, a bunch of them failing. I guess the failures were spurious given that my PR was merged without any reference to those.\nCode formatting #No particular code style seemed to be enforced for this bash script (no idea if there actually are bash formatters out there). I just imitated the surrounding style.\nGovernance and roadmap #According to the README, this Docker image is maintained by the community and not Nextcloud GmbH, the company behind Nextcloud. Maybe I should consider migrating to the other one (nextcloud/all-in-one), but given its description it looked like it would be a lot heavier. Who knows!\nWould I contribute again? #Only if I really have to!\n","date":"9 January 2024","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-nextclouds-docker-image/","section":"Posts","summary":"\u003cp\u003eWelcome to my second contribution experience report. It follows\n\u003ca href=\"/posts/contribution-experience-report-git\"\u003ethe one for Git\u003c/a\u003e which gives some background about the initiative.\u003c/p\u003e\n\u003ch3 id=\"my-motivation-to-contribute\" class=\"relative group\"\u003eMy motivation to contribute \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#my-motivation-to-contribute\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eIn short: I had been hit by a very annoying bug and I wanted to fix it.\u003c/p\u003e\n\u003cp\u003eHere is more background. In \u003ca href=\"https://kanthaus.online\" target=\"_blank\" rel=\"noreferrer\"\u003eKanthaus\u003c/a\u003e we have a NextCloud instance that we\nuse for shared documents. I have recently started helping out to maintain it,\nwhich means doing regular upgrades.\u003c/p\u003e","title":"Contribution experience report: NextCloud's Docker image"},{"content":"Improving the experience of new contributors in OpenRefine or other projects I maintain is an important topic for me. I think that whether a contributor stays active in a project depends a lot on the experience they have during their first contact. I would generally like to have more feedback about the hurdles people have when trying to contribute to projects where I am active. So I am starting to document my own experience when contributing to other projects. This feels like a useful way to take notes of things I would like to take inspiration from. Perhaps it is interesting to the said projects too. Of course, I don\u0026rsquo;t claim that this experience is representative. Also, I acknowledge that not every open source project is necessarily seeking new contributors: it can be a deliberate choice not to invest any energy in onboarding people or to refuse any external contributions, for various reasons. It\u0026rsquo;s not what I wish for the projects I am involved in, though.\nLet\u0026rsquo;s start with a first report about contributing to the Git project.\nMy motivation to contribute #In short:\nI wanted to implement a feature I needed, which I thought could also be useful to others it felt like an exciting challenge to submit a contribution to this venerable project Here is the longer version. I currently spend quite some time merging or rebasing things in OpenRefine and I have thought it would be wise to invest in a bit of tooling to ease that. I have discovered that git makes it possible to define custom merge drivers. This lets users change the algorithm used to merge two diverging versions of a given file together. You can provide your own executable file whose task is, given the two diverging versions and the common base version of a file, compute the merged file (possibly with merge conflict markers in it).\nWhen writing my own merge driver, I wanted to use git\u0026rsquo;s own algorithms as a starting point, and thankfully this is possible via the git merge-file command, which is essentially a command-line interface to the natively available merge drivers in git. However, I quickly noticed that this interface would sometimes give worse results than what I would get when letting git merge use its default merge driver. This is because the newer algorithms developed for git merge hadn\u0026rsquo;t been made available to the lesser known git merge-file command.\nSo it felt like an exciting opportunity to try and make it possible to use those newer algorithms in git merge-file. Knowing the lineage and success of the git project, it felt quite daunting but also very exciting to submit a contribution there. That added to the motivation.\nFirst contact with the project #Technically this wasn\u0026rsquo;t my first contact with the project as a contributor: I can\u0026rsquo;t resist bragging about the fact that I had already got a commit in 7 years earlier. Granted, it was only adding a single letter to a translation file, but you know, it\u0026rsquo;s still a commit.\nAlthough there is an official git project on GitHub, it\u0026rsquo;s not the place where the discussions take place. The real deal is the pretty scary official mailing list, where everything happens: project discussion, but also code review because contributions are sent as patches directly on the mailing list.\nI found the mailing list \u0026ldquo;scary\u0026rdquo; because it looked like a big forest of patches, mentioning a lot of things I never heard of. I was vaguely aware that people have pretty specific customs about how to behave and how to format messages there. I was pretty sure I was going to fail observing those traditions, outing me as a clueless noob, or maybe even an annoying spammer if I messed up really bad.\nSo because it looked a bit daunting, I didn\u0026rsquo;t even try to discuss my ideas with the project before implementing them. I think it\u0026rsquo;s generally much better to take the time to do discuss before, because it can save you a lot of time if the maintainers disagree with your approach. But given that my change felt pretty straightforward, I wanted to try to just craft the perfect patch directly. It would be just so clearly right and good that it would get waived through the reviewing pipeline and released to the world. Wooooosh!\nDevelopment environment #For me it\u0026rsquo;s usually a bit of a chore to install and get used to a new development environment required by a project. But in this case, it was like a visit to an adorable curiosity cabinet.\nGit itself has so few dependencies that I already had basically everything I needed on my Debian machine. I ran make and interatively installed some missing libraries or tools via APT. That was it, it compiled. To make changes to the code, I just used vanilla Vim.\nNo need for the latest version of that fancy package manager that you first need to install by piping a curl into a sudo. No need for virtual environments, long Docker pulls, or anything like that. It felt like a different universe. Surprisingly slick and enjoyable.\nFinding my way into the code base #That was a little trickier. I used my standard technique of searching through all files in the repository with grep, gradually backtracking from the command line options to the internal data structures used to represent the configuration of diff algorithms, through the various interfaces. I did write C code in the past, so I wasn\u0026rsquo;t completely new to its oddities, but still, trying to figure out how to properly use a macro or how to set a flag with bitwise operators felt like quite an adventure.\nTo continue with the curiosity cabinet: a lot of the source files are just at the root of the repository, not even in a folder. Being now used to deep folder structures from the Java world, I find that hilarious (in an enjoyable way). Also, I thought it was the norm to use preprocessor statements to make include guards preventing double includes, but apparently not in this code base. I had no idea it was manageable to do without for a project of that scale. (There are indeed include guards, not sure how I missed them. Thanks to Andrew Clayton for pointing out my mistake.)\nReviewing experience #To submit my changes, I could have tried to use git send-email, but it felt much easier to make a pull request to the GitHub repository as that\u0026rsquo;s a workflow I am used to. The GitGitGadget tool is a sort of GitHub bot that takes care of translating such pull requests to patches of the expected format and send those to the mailing list. For my previous pull request, I had used submitGit, a similar tool that was in use at the time (it does not seem to work anymore).\nThe timeline of my patch was as follows:\nNov. 8: submitted the first version of the patch, without tests (as a way to ask if my approach was the right one). Nov. 17: not having had any reaction, I wrote a reply to the patch, expanding on my motivation and other approaches I considered. Nov. 19: first reply from Philipp Wood, with detailed comments about my changes. It felt supportive and encouraging. Nov. 20: reply from the maintainer Junio Hamano, which also looked supportive. I submitted a new version of my patch, with tests, on the same day. Nov. 21: \u0026ldquo;looks good to me\u0026rdquo; reply from Philipp Wood. Nov. 23: patch merged in the \u0026ldquo;seen\u0026rdquo; branch. Dec. 11: patch merged in the \u0026ldquo;next\u0026rdquo; branch. Dec. 19: patch merged in the \u0026ldquo;master\u0026rdquo; branch. The communication during the review process was very friendly and helpful. At some point, I added the metadata \u0026ldquo;Reviewed-by: Philipp Wood\u0026rdquo; (with email address) to my commit with the intention of acknowledging his reviewing effort, but it turned out that it was a faux-pas: I guess it probably implies that the person approved the changes, while at this point he had not. This breach of the etiquette was pointed out to me in a most excellent way.\nTesting infrastructure #Git\u0026rsquo;s test suite is written as a collection of bash scripts which test it via its command-line interface. Given that Git\u0026rsquo;s command-line interface is its canonical user interface, those are de facto end-to-end tests. There are ongoing efforts to introduce unit tests written in C, but it looks like this is still not ready for contributors to adopt.\nTests being bash scripts, they can just be run as such, which is again quite convenient: no need to learn a testing framework, just run the bash file that contains the test you care about.\nOn top of that, the GitHub repository comes with a collection of pull request checks which run this test suite and a lot of other things. That was also super convenient to make sure I did not break anything with my changes, without having to learn how to invoke those checks on my own (and without having to provide the computation resources for it). I have no idea how the maintainers work with this test suite since they are not using GitHub (or any other forge, as far as I know), but as an external contributor, it is useful.\nCode formatting #My editor inserted spaces instead of tabs for indentation, which conflicted with the default style. I think this was caught either by git itself, by highlighting the whitespace in the diff view, or by the pull requests checks on GitHub. The pull request checks might also have brought up other issues that I don\u0026rsquo;t remember: in any case, it was simple enough to fix those as as a reaction to that.\nGovernance and roadmap #As I understand it, this project is structured around a central maintainer, Junio C Hamano, who has the final responsibility for accepting patches and leading the release process. Other contributors (such as Philipp Wood in my case) seem to help with reviewing, but I don\u0026rsquo;t know to what extent this role is formalized. The current maintainer seems to have been designated by Linus Torvalds directly, so I guess the understanding is that he remains the only one ultimately in charge until he designates someone else. This is all guesswork on my part without having looked into it at all, just based on my existing interactions with the project. This rather blurry governance (from my perspective) was not an obstacle to me, beyond the need to find someone to review my changes in the first place (which happened by itself).\nI had a look at the minutes from the Git Contributor\u0026rsquo;s Summit 2023 which I found very interesting. They give a sense of where the project is heading to and that people are interested in easing the onboarding of newcomers.\nWould I contribute again? #Absolutely. It would likely be focused on something I need myself though - I don\u0026rsquo;t really see myself contributing to git for its own sake, because I think there are a lot of people in big tech companies who are in a better position to do so and I think it\u0026rsquo;s right that those companies contribute back.\n","date":"17 December 2023","permalink":"https://antonin.delpeuch.eu/posts/contribution-experience-report-git/","section":"Posts","summary":"\u003cp\u003eImproving the experience of new contributors in \u003ca href=\"https://openrefine.org\" target=\"_blank\" rel=\"noreferrer\"\u003eOpenRefine\u003c/a\u003e or other projects I maintain is an important topic for me.\nI think that whether a contributor stays active in a project depends a lot on the experience they have during their first contact.\nI would generally like to have more feedback about the hurdles people have when trying to contribute to projects where I am active.\nSo I am starting to document my own experience when contributing to other projects. This feels like a useful way to take notes of things I would like to take inspiration from. Perhaps it is interesting to the said projects too.\nOf course, I don\u0026rsquo;t claim that this experience is representative. Also, I acknowledge that not every open source project is necessarily seeking new contributors: it can be a deliberate choice not to invest any energy in onboarding people or to refuse any external contributions, for various reasons. It\u0026rsquo;s not what I wish for the projects I am involved in, though.\u003c/p\u003e","title":"Contribution experience report: Git"},{"content":"In OpenRefine, we have been enforcing a code style for our Java files using a linter, which reformats source files according to a configuration expressed in Eclipse\u0026rsquo;s internal format.\nBecause the linter reuses Eclipse\u0026rsquo;s internal libraries, it is of course written in Java. Also, we invoke it via Maven, meaning that the start up time of the linter is quite long: we need to boot a Java virtual machine (JVM), which then boots Maven, which finally runs the formatter. So that\u0026rsquo;s not exactly fast. On my laptop it takes about 11 seconds. It\u0026rsquo;s not the end of the world, but it means it\u0026rsquo;s a bit annoying to have it as a git pre-commit hook, and it\u0026rsquo;s definitely not suitable as a git filter driver. Those filter drivers are processes which take the contents of a file on standard input and output a normalized version of it. Given that they are run on the fly for things as basic as doing a git diff or git status, they need to be really quick.\nGraalVM #For a while I have been looking at Oracle\u0026rsquo;s GraalVM project, mostly with an interest in the inter-language interoperability and interpreter optimization features because it could be useful for OpenRefine\u0026rsquo;s integration of Python and other expression languages. I won\u0026rsquo;t reproduce the sales pitch here, there are plenty of resources out there which explain it much better than I\u0026rsquo;d be able to. But it seems that the most popular feature of this technology is a fairly different one: the ability to generate \u0026ldquo;native images\u0026rdquo;, meaning the compilation of a Java program to native code. By doing so it makes it possible to boot the program much faster, so it seems to be useful in cloud environments where processes are started up on demand, for instance when an HTTP request is coming in.\nSo this felt like a good opportunity to try out this technology: can we turn this linter into a native image, such that it\u0026rsquo;s fast enough to be run as a git filter?\nPutting something together #As a quick experiment, I started out a small Java program to lint Java files according to OpenRefine\u0026rsquo;s style. It\u0026rsquo;s basically doing the same thing as the Maven plugin we use, but without going through Maven at all, so that we can at least directly avoid this slowness. It\u0026rsquo;s calling Eclipse\u0026rsquo;s own formatting code using a hard-coded configuration that matches OpenRefine\u0026rsquo;s settings. Because we also started normalizing our import order, I went ahead and added calls to another library which does that (that library is actually a Maven plugin, but luckily we can avoid pulling in the Maven dependencies by excluding them explicitly, and only call Maven-free code).\nWhen running this Java program normally, via a .jar file which includes all necessary dependenciesi, it takes about 0.8 second to format a fairly big Java file. That\u0026rsquo;s still too long for a filter driver.\nTurning the program into a binary #To turn this program into a binary, we first need to download GraalVM, which is a sort of alternative Java distribution with all those fancy features enabled. There is an open source version, called the \u0026ldquo;Community Edition\u0026rdquo; (released under LGPL v2). Once we are running that, there is a Maven plugin which helps generating the native image as part of Maven\u0026rsquo;s build process, which is quite helpful. It even supports compiling and running the Java tests as a native binary. To generate the native image (skipping the tests), one can just run:\nmvn -Pnative -DskipTests -Dagent package Well, \u0026ldquo;just\u0026rdquo; is a big word because this process is actually really intense: the process takes in total 2m 22s to compile my little linter (on a more beefy desktop), which is turned into a 61 MB binary. We get an overview of the compilation phases:\n[1/8] Initializing... (4,1s @ 0,17GB) Java version: 21.0.1+12, vendor version: Oracle GraalVM 21.0.1+12.1 Graal compiler: optimization level: 2, target machine: x86-64-v3, PGO: ML-inferred C compiler: gcc (linux, x86_64, 13.2.0) Garbage collector: Serial GC (max heap size: 80% of RAM) 1 user-specific feature(s): - com.oracle.svm.thirdparty.gson.GsonFeature ----------------------------------------------------------------------------------------------------------------------- Build resources: - 11,75GB of memory (75,6% of 15,54GB system memory, determined at start) - 8 thread(s) (100,0% of 8 available processor(s), determined at start) [2/8] Performing analysis... [***] (34,0s @ 0,95GB) 9 142 reachable types (83,5% of 10 948 total) 16 427 reachable fields (61,2% of 26 857 total) 55 394 reachable methods (61,8% of 89 703 total) 2 749 types, 466 fields, and 1 313 methods registered for reflection 59 types, 56 fields, and 53 methods registered for JNI access 4 native libraries: dl, pthread, rt, z [3/8] Building universe... (4,0s @ 1,43GB) [4/8] Parsing methods... [****] (13,1s @ 1,51GB) [5/8] Inlining methods... [****] (1,7s @ 1,89GB) [6/8] Compiling methods... [********] (75,4s @ 1,92GB) [7/8] Layouting methods... [***] (5,8s @ 1,25GB) [8/8] Creating image... [**] (3,4s @ 1,69GB) 36,34MB (59,58%) for code area: 33 145 compilation units 22,27MB (36,50%) for image heap: 249 379 objects and 85 resources 2,39MB ( 3,92%) for other data 60,99MB in total Now let\u0026rsquo;s start it!\n$ ./target/filterlinter \u0026lt; Example.java Exception in thread \u0026#34;main\u0026#34; java.lang.ExceptionInInitializerError at java.base@21.0.1/java.lang.Class.ensureInitialized(DynamicHub.java:595) at org.eclipse.jdt.core.dom.CompilationUnitResolver.parse(CompilationUnitResolver.java:604) at org.eclipse.jdt.core.dom.ASTParser.internalCreateAST(ASTParser.java:1264) at org.eclipse.jdt.core.dom.ASTParser.createAST(ASTParser.java:868) at org.eclipse.jdt.internal.formatter.DefaultCodeFormatter.parseSourceCode(DefaultCodeFormatter.java:317) at org.eclipse.jdt.internal.formatter.DefaultCodeFormatter.prepareFormattedCode(DefaultCodeFormatter.java:221) at org.eclipse.jdt.internal.formatter.DefaultCodeFormatter.format(DefaultCodeFormatter.java:185) at eu.delpeuch.antonin.filterlinter.Formatter.format(Formatter.java:49) at eu.delpeuch.antonin.filterlinter.App.formatString(App.java:41) at eu.delpeuch.antonin.filterlinter.App.main(App.java:35) at java.base@21.0.1/java.lang.invoke.LambdaForm$DMH/sa346b79c.invokeStaticInit(LambdaForm$DMH) Caused by: java.lang.NullPointerException at java.base@21.0.1/java.text.MessageFormat.applyPattern(MessageFormat.java:468) at java.base@21.0.1/java.text.MessageFormat.\u0026lt;init\u0026gt;(MessageFormat.java:382) at java.base@21.0.1/java.text.MessageFormat.format(MessageFormat.java:882) at org.eclipse.jdt.internal.compiler.util.Messages.bind(Messages.java:173) at org.eclipse.jdt.internal.compiler.util.Messages.bind(Messages.java:150) at org.eclipse.jdt.internal.compiler.parser.Parser.readTable(Parser.java:816) at org.eclipse.jdt.internal.compiler.parser.Parser.initTables(Parser.java:658) at org.eclipse.jdt.internal.compiler.parser.Parser.\u0026lt;clinit\u0026gt;(Parser.java:175) ... 11 more Oops, that does not look like the linted Java code I wanted!\nThe catch #After searching the web to understand what this could be due to, I realized that it\u0026rsquo;s because the GraalVM compiler needs a bit of help to handle some corner cases. There are aspects of the Java language that it cannot simply compile by itself, such as the use of reflection. Reflection is the dynamic inspection of Java classes by the Java code itself, which can be used to do all sorts of funky things. For this to work in the native image, you basically need to tell the compiler ahead of time which classes will be inspected, so that it can pre-compute and store the output of those introspection calls.\nGraalVM offers an \u0026ldquo;agent\u0026rdquo; to solve that problem. You can run the original Java program with this agent enabled and it will record which of those reflection calls are made (and other sorts of special calls). That generates a set of configuration files to be used by the compiler, which in my case looked like that:\n[ { \u0026#34;name\u0026#34;: \u0026#34;com.github.javaparser.ast.body.FieldDeclaration\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;: true }, { \u0026#34;name\u0026#34;: \u0026#34;com.github.javaparser.ast.expr.VariableDeclarationExpr\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;: true }, { \u0026#34;name\u0026#34;: \u0026#34;java.util.concurrent.ForkJoinTask\u0026#34;, \u0026#34;fields\u0026#34;: [ { \u0026#34;name\u0026#34;: \u0026#34;aux\u0026#34; }, { \u0026#34;name\u0026#34;: \u0026#34;status\u0026#34; } ] }, { \u0026#34;name\u0026#34;: \u0026#34;java.util.concurrent.atomic.AtomicBoolean\u0026#34;, \u0026#34;fields\u0026#34;: [ { \u0026#34;name\u0026#34;: \u0026#34;value\u0026#34; } ] }, { \u0026#34;name\u0026#34;: \u0026#34;jdk.internal.misc.Unsafe\u0026#34; }, { \u0026#34;name\u0026#34;: \u0026#34;org.eclipse.core.internal.runtime.Messages\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;: true }, { \u0026#34;name\u0026#34;: \u0026#34;org.eclipse.jdt.internal.compiler.util.Messages\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;: true } ] Great! If we compile again and run the binary on the same file we run the agent on, we do get linted code out this time. And it\u0026rsquo;s indeed faster: 17 ms total time for a small file that would take 375 ms if we use the .jar instead. Nice.\nBut if I try running the linter again on a new file… I get a new error!\n$ ./target/filterlinter \u0026lt; Example2.java Exception in thread \u0026#34;main\u0026#34; java.lang.NoSuchFieldError: levels at com.github.javaparser.metamodel.PropertyMetaModel.getValue(PropertyMetaModel.java:260) at com.github.javaparser.ast.validator.language_level_validations.chunks.CommonValidators.lambda$new$7(CommonValidators.java:63) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:38) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.TreeVisitorValidator.accept(TreeVisitorValidator.java:40) at com.github.javaparser.ast.validator.Validators.lambda$accept$0(Validators.java:64) at java.base@21.0.1/java.util.ArrayList.forEach(ArrayList.java:1596) at com.github.javaparser.ast.validator.Validators.accept(Validators.java:64) at com.github.javaparser.ast.validator.Validators.lambda$accept$0(Validators.java:64) at java.base@21.0.1/java.util.ArrayList.forEach(ArrayList.java:1596) at com.github.javaparser.ast.validator.Validators.accept(Validators.java:64) at com.github.javaparser.ParserConfiguration$2.postProcess(ParserConfiguration.java:308) at com.github.javaparser.JavaParser.parse(JavaParser.java:128) at com.github.javaparser.JavaParser.parse(JavaParser.java:305) at net.revelc.code.impsort.ImpSort.parseFile(ImpSort.java:129) at eu.delpeuch.antonin.filterlinter.Formatter.sortImports(Formatter.java:67) at eu.delpeuch.antonin.filterlinter.App.formatString(App.java:43) at eu.delpeuch.antonin.filterlinter.App.formatFile(App.java:53) at eu.delpeuch.antonin.filterlinter.App.main(App.java:30) at java.base@21.0.1/java.lang.invoke.LambdaForm$DMH/sa346b79c.invokeStaticInit(LambdaForm$DMH) That\u0026rsquo;s where I think the technology starts to get much less convincing. The problem here, as the stack trace might suggest, is that the Java parser we rely on uses reflection internally. By running the agent on an example file, the uses of reflection were only captured for the syntactic constructs that were present in the example file: in the JSON configuration above, you see for instance a mention of com.github.javaparser.ast.body.FieldDeclaration, which must be a class that represents Java field declarations in an abstract syntax tree. So whenever we run our binary on another file with unseen constructs, we are missing the required information to simulate the reflection calls, and we just fail.\nI really wonder what the expected workaround is for this sort of situation. Of course, it is natural to try and execute the agent on a set of files that\u0026rsquo;s as diverse as possible, but I don\u0026rsquo;t want the correctness of my program to rely on me finding a set of Java files which cover the entire set of possible AST nodes! Even if my code base had Java tests featuring 100% code coverage, executing my tests with the agent would not be enough, since those reflection calls are done by a dependency, not my own code.\nBecause this was just an experiment, I decided to go for a hacky route. Just manually add all the AST nodes to the JSON file. I can first generate the list of AST nodes by inspecting the .jar file:\n$ jar -tf target/filterlinter-0.0.1-SNAPSHOT-jar-with-dependencies.jar| grep com.github.javaparser.ast | grep -P \u0026#34;\\.class$\u0026#34; com/github/javaparser/ast/AccessSpecifier.class com/github/javaparser/ast/AllFieldsConstructor.class com/github/javaparser/ast/ArrayCreationLevel.class com/github/javaparser/ast/body/AnnotationDeclaration.class com/github/javaparser/ast/body/AnnotationMemberDeclaration.class com/github/javaparser/ast/body/BodyDeclaration.class ... That\u0026rsquo;s 282 classes in total. Then with a bit more bash scripting I can translate this list into the required JSON objects:\n[ { \u0026#34;name\u0026#34;:\u0026#34;com.github.javaparser.ast.AccessSpecifier\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;:true }, { \u0026#34;name\u0026#34;:\u0026#34;com.github.javaparser.ast.AllFieldsconstructor\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;:true }, { \u0026#34;name\u0026#34;:\u0026#34;com.github.javaparser.ast.ArrayCreationLevel\u0026#34;, \u0026#34;allDeclaredFields\u0026#34;:true }, ... And recompile my native image with this new configuration. That\u0026rsquo;s incredibly ugly, but it seems to work, and it still runs fast.\nOf course it would be much easier if you could at least ask for all the classes in a given Java package to have reflection enabled. I was pleased to add the 42nd thumbs up on this GitHub issue. But even if that was possible, I really wonder what is the expected solution for this problem. How can I be really sure there isn\u0026rsquo;t another reachable use of reflection somewhere, that I just haven\u0026rsquo;t encountered yet?\nNeedless to say, the resulting binary is not something I can propose to use as a linter in the OpenRefine project: I would need to distribute one for all reasonable operating systems and CPU architectures and embedding even a single 61 MB binary in the git repository is not going to fly. It was fun to try this out regardless, and I might still use the binary myself, as it is actually fast enough to be run as a filter in my opinion.\n","date":"16 November 2023","permalink":"https://antonin.delpeuch.eu/posts/trying-out-graals-native-image-functionality/","section":"Posts","summary":"\u003cp\u003eIn \u003ca href=\"https://openrefine.org\" target=\"_blank\" rel=\"noreferrer\"\u003eOpenRefine\u003c/a\u003e, we have been enforcing a code style for our Java files using \u003ca href=\"https://github.com/revelc/formatter-maven-plugin\" target=\"_blank\" rel=\"noreferrer\"\u003ea linter\u003c/a\u003e,\nwhich reformats source files according to a configuration expressed in Eclipse\u0026rsquo;s internal format.\u003c/p\u003e\n\u003cp\u003eBecause the linter reuses Eclipse\u0026rsquo;s internal libraries, it is of course written in Java. Also, we invoke it via Maven, meaning that\nthe start up time of the linter is quite long: we need to boot a Java virtual machine (JVM), which then boots Maven, which finally runs\nthe formatter. So that\u0026rsquo;s not exactly fast. On my laptop it takes about 11 seconds. It\u0026rsquo;s not the end of the world, but it means it\u0026rsquo;s\na bit annoying to have it as a git \u003ca href=\"https://git-scm.com/book/en/v2/Customizing-Git-Git-Hooks\" target=\"_blank\" rel=\"noreferrer\"\u003epre-commit hook\u003c/a\u003e, and it\u0026rsquo;s definitely\nnot suitable as a git \u003ca href=\"https://git-scm.com/docs/gitattributes#_filter\" target=\"_blank\" rel=\"noreferrer\"\u003efilter driver\u003c/a\u003e. Those filter drivers are processes which take\nthe contents of a file on standard input and output a normalized version of it. Given that they are run on the fly for things as basic\nas doing a \u003ccode\u003egit diff\u003c/code\u003e or \u003ccode\u003egit status\u003c/code\u003e, they need to be really quick.\u003c/p\u003e","title":"Trying out Graal's Native image functionality"},{"content":"In OpenRefine we still haven\u0026rsquo;t got ourselves a clear, public roadmap indicating what we are working towards in the near future. There are various proposals as to how we should get ourselves one.\nSo I thought I would at least write down what my own priorities are and paint of my vision for the project. I don\u0026rsquo;t claim that this should be OpenRefine\u0026rsquo;s roadmap: there is a ton of other efforts that I would really agree with, but which I am less interested in working on myself. Other team members have their own agenda and my own priorities are not set in stone either.\nHere is a short summary of my personal goals:\nGrow the team and set up the project so that it runs sustainably Reproducibility Reconciliation Extensibility Goal 1: Grow the team and set up the project so that it runs sustainably #This goal is not about which features OpenRefine should have and isn\u0026rsquo;t about software development per se, so maybe that\u0026rsquo;s not what you expected to read first. But that\u0026rsquo;s something I dedicate a good chunk of my time to.\nThe bottom line is that I won\u0026rsquo;t be around forever in this project and I would be really happy to leave it in a state where I would be reasonably confident that an enthusiastic, talented and friendly team would keep running for a while after my departure.\nAlthough I have this as a personal goal, it\u0026rsquo;s also clear I am doing a lot of things wrong in this regard. I have definitely been learning quite a lot on this topic for the past few years, also because of my involvement in Kanthaus, a house project which shares similar challenges, with the constant need to onboard more people to keep the community active and dynamic.\nGoal 2: Reproducibility #One of the things that I really liked when I first discovered OpenRefine was the ability to extract the history and replay it on a new version of the dataset. It\u0026rsquo;s such a great idea and could be potentially so useful! Sadly it\u0026rsquo;s really far from reliable so I really want to improve that.\nFirst, why is this feature important to me?\nIt\u0026rsquo;s something a lot of people really need! The research community is an obvious example, given the realization that a lot of scientific results cannot be reliably replicated based on the published description of the research methods. So, if someone uses OpenRefine at some point in their research workflow, I want it to be really easy to inspect their cleaning data process and re-run it in a new environment, on a slightly different dataset. But it\u0026rsquo;s not just about research! OpenRefine is also used a lot to import data in knowledge graphs, such as Wikidata. Importing a dataset in a knowledge graph is rarely a one-off task, because those datasets often get updated. If we can reliably re-run data cleaning processes made with OpenRefine, then we are really close to having a clean infrastructure for continuously importing datasets in Wikidata. You could have an online platform which hosts those workflows, making it possible for other contributors to audit and update them as the target data model evolves. Many of the architectural changes that are required for this reproducibility also enable other features, seemingly unrelated, which can really help users even if they do not care about re-running their data cleaning process on a new dataset later on. One of the most long-standing open issues we had in OpenRefine was about the grid view being reset to the first page whenever the user would run an operation that could affect multiple rows in the grid (such as matching a reconciled cell to a particular entity). Determining if the current view on the grid can meaningfully be preserved requires assigning suitable metadata to operations. This is the same sort of metadata that is required to determine if a series of operations can be applied on a new grid, or if two operations can safely be run in parallel. I have the impression that for many users, OpenRefine is the first step in their journey towards more principled data transformations. They might have done manual edits in Excel before and might move on to Jupyter notebooks later. I want to believe that OpenRefine can be a meaningful step on this data literacy path, in that it still lets the user carry out transformations interactively, via a user interface, but introduces them to some concepts of programming. The fact that operations run on all rows by default. The combining of filters to select a set of rows. The gradual exposure to regular expressions or expression languages such as GREL or Python. And of course, this ability to re-run the history on a new file, which means that the user is, in some way, interactively writing a small program without thinking about it. But that can only be convincing if this last feature really lives up to the legitimate expectations that users develop about it - which is not the case yet. I am currently working on this very topic (in the scope of a funded project) and you can see an overview of the progress in this post. In the Development \u0026amp; Design category of OpenRefine\u0026rsquo;s forum I also post other updates, often with screencasts to demonstrate the features I am working on.\nGoal 3: Reconciliation #Reconciliation is another feature of OpenRefine which got me hooked when I first discovered the tool. It struck me as a really well designed and approachable solution to such an important problem. I was looking for a tool to match a dataset against Wikidata and there was really not a lot of decent options. Magnus Manske\u0026rsquo;s Mix\u0026rsquo;n\u0026rsquo;Match tool was a helpful attempt to crowd-source matching of various databases with Wikidata, but made it hard to finely tune the matching or automate parts of it and its integration with other steps of a data import process was fairly unclear. Magnus had also written a reconciliation service for Wikidata, which was not working very well in OpenRefine. It\u0026rsquo;s hard to blame him for that: the protocol was pretty poorly documented and one essentially had to inspect OpenRefine\u0026rsquo;s own source code to understand what the server was supposed to do. So that\u0026rsquo;s what led me to writing a Wikidata reconciliation service and then to get involved in OpenRefine\u0026rsquo;s development.\nSo what is there left to do? Many things!\nAlthough I was initially enthusiastic about the design and user-friendliness of the feature, there are tons of things to improve. Reconciliation is a very tricky task and as a tool developer I feel like I have some duties which I don\u0026rsquo;t really fulfill yet in this area:\na duty not to deceive users. That means conveying the correct expectations about the matches suggested by reconciliation services and encourage them to review those critically a duty not to waste people\u0026rsquo;s time. That means giving them time-efficient and intuitive workflows to configure reconciliation, review and improve its results. Having fast reconciliation services is not sufficient for that: the time I primarily want to minimize is not so much the time spent waiting for reconciliation results to come in, but rather the time spent reviewing and correcting reconciliation results. a duty to encourage sound methodologies. We don\u0026rsquo;t do a good job of directing users towards a scientifically valid workflow (train and tune reconciliation on a sample of a dataset, evaluate its performance on another sample) which is the only way they can get a sense of what accuracy they are achieving. Behind the scenes, the reconciliation protocol is still something that was generalized from Freebase\u0026rsquo;s own APIs and the current specifications still bear traces of that, with oddities which don\u0026rsquo;t make so much sense outside of Freebase, or things that simply can be improved more than 10 years later, in a world where people\u0026rsquo;s expectations about APIs have changed a little. I think this protocol deserves to be adopted broadly, by a lot of data sources and a lot of clients wanting to do matching with those sources. That\u0026rsquo;s why we started a Community Group within W3C. It gathers a lot of people (48 to date) interested in the protocol for one reason or another, who come together to improve its specifications. A few years after the creation of the group, we have done quite a lot already: documenting the existing protocol (version 0.1), releasing a first improved version of the specs (version 0.2) and drafting a lot of changes for the next version. We have documented the ecosystem around the protocol, built a test bench to help developers check the compliance of their service with the specifications and new services and tools using the protocol have appeared. But a lot of those improvements still haven\u0026rsquo;t reached OpenRefine users so that\u0026rsquo;s something I want us to catch up on.\nConcretely, what do I want to get done? We (meaning Ayushi, Lydia, Lozana and I) are currently working (in the scope of a funded project) on various usability improvements for our existing reconciliation features. Beyond this effort, I would like to work on the following:\nupdating OpenRefine so it can take advantage of the newer features offered by the protocol, in its current draft. For instance, exposing the reconciliation features returned by the services to the user, so that they can rely on something better than an opaque matching score. base OpenRefine\u0026rsquo;s reconciliation features on an external, reusable pair of libraries (one backend-side in Java, another frontend-side in JS). I see this as an opportunity to do a big clean up in this code base and encourage the use of the protocol in other clients. The libraries would also help with supporting multiple versions of the protocol, as I would like that OpenRefine remains compatible with as many existing services as possible. I have started working on such a Java library (provisionally named ReconToolkit). Improve the tool to let the user easily build a robust set of criteria to match a column in their dataset to an external service. This could take the form of interactively asking the user to make judgments on certain rows, selected by an active learning algorithm, to construct a matching classifier fitted to the dataset at hand (based both on features exposed by the reconciliation service and other ones computed locally). Such a model could then be evaluated on another set of rows to get an estimate of its performance. It could be made reusable and shareable. Although I am phrasing those features in fairly technical terms, I suspect it might be possible to present this to users in an approachable way, even for those not familiar with machine learning or statistics. There are other things I would really like to see happen (and I would likely help with), related to reconciliation but outside the scope of OpenRefine itself, such as:\nGetting the reconciliation protocol to become a W3C recommendation. Not that I think the actual status makes a ton of difference for adoption, but rather that it would be a sign that we are happy enough about the state of the specs and their implementations that we apply for this. We have already made many very useful improvements to the specs by attempting to comply with W3C\u0026rsquo;s guidelines (for instance thanks to Fabian Steeg\u0026rsquo;s tireless work on internationalization matters), so I find this process quite helpful so far Having a nice, user-facing directory of reconciliation services. The list in the Test bench is not designed for that Having a Wikibase extension which implements the reconciliation protocol. I am hoping this would make reconciliation more reliable, faster and easier to install in a Wikibase instance Goal 4: Extensibility #Data cleaning needs are very diverse and there is no chance that OpenRefine fulfills enough of those out of the box. People import data from various sources, need various sorts of cleaning steps and export the results to various places. To cater for that, OpenRefine has an extension system, which lets third parties add features to OpenRefine without modifying it directly.\nAlthough we already have an extension system, it\u0026rsquo;s far from working as it should. First, the experience of installing an OpenRefine extension is quite technical, which likely puts off many users. There should be an easy way to install an extension from the application itself. The same goes of course for upgrading and uninstalling. The stability guarantees are very thin: it\u0026rsquo;s easy for an extension to crash the entire app, for instance if it was designed to work with a different version of OpenRefine. The development process of extensions is also poorly documented. All this means that the hurdles for a third party to start developing and then keep maintaining an extension are very high.\nI think it\u0026rsquo;s key to the sustainability of the project that this extension system works well. There is a lot of potential in custom integrations with specific platforms, to help people ingest data following a specific data model. Our Wikibase integration is used a lot and the community is asking us for more. The RDF Transform extension is also a popular one. There are extensions to work with OpenStreetMap data too. The Ontotext Refine fork, developed to help people ingest data in the OntoDB triple store, recently got promoted to a standalone product and they made it possible to install OpenRefine extensions in it. We cannot have built in support for all those platforms out of the box, so it\u0026rsquo;s important that other teams are able to take responsibility for this development.\nBecause OpenRefine has a fairly uncommon architecture, by being a web app that people run locally, there aren\u0026rsquo;t a lot of existing application platforms to pick from. OpenRefine extensions must be able to add new functionality both server-side and client-side, in the same extension. I am not aware of any framework which offers that. We have been discussing possible strategies on OpenRefine\u0026rsquo;s forum for a while now and I feel like we are slowly getting somewhere, but there are still a lot of open questions in my opinion.\n","date":"21 October 2023","permalink":"https://antonin.delpeuch.eu/posts/my-roadmap-for-openrefine/","section":"Posts","summary":"\u003cp\u003eIn \u003ca href=\"https://openrefine.org\" target=\"_blank\" rel=\"noreferrer\"\u003eOpenRefine\u003c/a\u003e we still haven\u0026rsquo;t got ourselves a clear, public roadmap indicating what we are working towards in the near future.\nThere are \u003ca href=\"https://forum.openrefine.org/t/openrefine-2032-what-direction-does-openrefine-want-to-go/272\" target=\"_blank\" rel=\"noreferrer\"\u003evarious proposals\u003c/a\u003e as to how we should get ourselves one.\u003c/p\u003e\n\u003cp\u003eSo I thought I would at least write down what my own priorities are and paint of my vision for the project.\nI don\u0026rsquo;t claim that this should be OpenRefine\u0026rsquo;s roadmap: there is a ton of other efforts that I would really agree with, but which I am less interested in working on myself. Other team members have their own agenda and my own priorities are not set in stone either.\u003c/p\u003e","title":"My roadmap for OpenRefine"},{"content":"A lot of my work happens on software forges like GitHub or GitLab. They are very useful platforms, but they are not perfect. In this post I want to describe a few missing features around the transmission of open source projects to new maintainers.\nDiscovering active forks of a repository #A lot of projects are maintained by a single person who, after a while, loses interest in the project, or even moves on completely and stops responding to any activity on the repository.\nIn theory, that\u0026rsquo;s not a problem, right? If there is still a community of people relying on this project, they can \u0026ldquo;just fork it\u0026rdquo; and continue their activity there. Of course it\u0026rsquo;s not that easy: the project will often be associated with assets that cannot be reclaimed, like a domain name, a package name in some package repository, social media or crowdfunding accounts, and so on. But even ignoring all of those, one basic thing will need to change: the address of the repository.\nAnd this is where I find existing software forges quite disappointing. Most of them (like GitHub or GitLab) have a notion of \u0026ldquo;fork\u0026rdquo;, meaning that it\u0026rsquo;s possible to create a personal clone of a repository under one\u0026rsquo;s account (or as an organization), but this way of forking a project is rarely helpful when someone (or a group of people) actually want to continue development on a project because they are unable to do so in the original repository.\nWhy is pressing the \u0026ldquo;fork\u0026rdquo; buttons on GitHub or GitLab not a solution in such a situation?\nBecause they aren\u0026rsquo;t really meant to fork anything! Those so-called \u0026ldquo;forks\u0026rdquo; are primarily designed for a very different situation: making a contribution to a repository that is actively maintained. You fork the repository, create a branch there, make your changes and submit a pull request. One sign of this intention is that commits you make in your fork will not appear in the activity timeline on your profile: those will only be recognized as contributions once they are sent upstream.\nThe bigger problem, I think, is that people landing on the original repository will have a hard time discovering your actively maintained fork. GitHub\u0026rsquo;s \u0026ldquo;Network\u0026rdquo; page does make it possible to see the list of forks and get a rough sense of which ones seem to be active, but one needs to be quite motivated to go there and investigate the state of the project\u0026rsquo;s ecosystem in this way.\nWhat I would find useful is that when repositories lose the attention of their owners, the forge starts advertising the most active fork(s) on the main page of the repository. This would require some safeguards, as it would be tempting for spammers or crooks to exploit such a mechanism by creating artificially active forks of very popular repositories, so that they redirect the attention to them.\nSo how about taking some inspiration from those pedals train drivers need to regularly press to confirm they are still conscious? After one year (say) of the repository owners not being active in the repository, they are sent an email asking if they still feel responsible for the maintenance of that repository. If they do not confirm that they have it on their radar, the repository is marked as inactive and forks are displayed on its main page, ranked by some activity or popularity measure.\nI think a feature like this would really help avoiding discontinuing software projects needlessly. GitHub has this helpful \u0026ldquo;successor\u0026rdquo; feature, making it possible to transfer the ownership of your repositories to someone else when you die, but that\u0026rsquo;s of course an extreme case: a lot of people abandon their repositories well before they die. The conditions for this mechanism to kick in are really hard to meet: the owner must have known of the feature and used it before they died, then GitHub must get some sort of official confirmation of the death from some authorities (I suppose, at least) and finally the successor should be available to respond and take on the ownership or transfer it further.\nGranting more rights to new contributors #Consider the different situation of a healthy project, actively maintained.\nWhen a new contributor arrives in the project, there is generally no clear route for their onboarding. They might have made dozens of excellent contributions and have shown their commitment to the project over many months, yet they will generally not be allowed to do basic things such as labeling or closing issues. There is generally no clear message as to what bar needs to be met to obtain commit rights.\nI believe a lot of projects would benefit from more proactively granting project access to trusted contributors, but the because the forges do not incentivize that at all, they maintain the culture of open source contributors being performers on a scene, delivering their great sauce to an audience of \u0026ldquo;stargazers\u0026rdquo;, to put it in GitHub\u0026rsquo;s terms.\nCompare this to how other platforms grant rights automatically after some thresholds are met. On many Wikipedia editions, after you have made a bunch of edits, you become \u0026ldquo;autoconfirmed\u0026rdquo;, meaning that your edits to pages are no longer flagged as needing review. On StackExchange, after having earned enough reputation points, you can edit questions or answers written by others, or have access to some review queues. I am not aware of anything comparable on software forges.\nNow obviously, automatically granting rights to people after they have met certain thersholds has security implications. People can introduce bugs, vulnerabilities, spam or other sorts of nasty things in projects. Writing software is arguably a bit more sensitive than an encyclopedia or a Q\u0026amp;A site (although, I am sure you can make a lot of nasty stuff on such platforms too). It does make sense to keep the repository owners in the loop, especially to grant higher permissions such as the ability to publish releases. Just like you don\u0026rsquo;t automatically become an admin on Wikipedia, no matter how many edits you do.\nSo here are a few measures I am thinking about, which could lower the barrier to onboard people:\nHaving a way to request access to a repository. It\u0026rsquo;s such a basic thing, but GitHub does not have that (GitLab does). Whenever I think I could help out on a GitHub project, I try to identify who is in charge by looking at recent repository activity, and then try to find an email address for that person. Sometimes it\u0026rsquo;s visible in the commit log, sometimes not, so I would then try to stalk the person via a search engine, or awkwardly open an issue about my offer to help. One needs to be pretty self-assured and motivated to go that far! Having the forge actively suggest to the maintainers granting more rights to recurrent contributors. I would love to get a message along the lines of \u0026ldquo;this is the third pull request you approved from this person. Why not invite them to the repository?\u0026rdquo; You could even imagine that for some projects it would be appropriate to automatically grant such accesses after some threshold is met. The thing is, as a maintainer this is something I really don\u0026rsquo;t have on my radar all the time: I try to keep the issue and pull request backlog under control, publish releases regularly, respond to security advisories, do some user support, and probably other things on top. It\u0026rsquo;s already a lot of things to keep track of, so the task of proactively inviting people in the project gets easily forgotten. When our own interest in running a project fades away, it\u0026rsquo;s often too late to find a successor, so I think it\u0026rsquo;s really worth inviting people over earlier, when the excitement and activity are still running high. Some open source projects are by design meant to be only run by a certain team and are explicitly closed to other contributors. Some public Git repositories are merely used as online storage space with versioning, for a single user. That\u0026rsquo;s fine too, but I don\u0026rsquo;t think that should be the default that forges draw us towards.\n","date":"20 October 2023","permalink":"https://antonin.delpeuch.eu/posts/some-missing-features-in-software-forges/","section":"Posts","summary":"\u003cp\u003eA lot of my work happens on software forges like GitHub or GitLab. They are very useful platforms, but they are not perfect. In this post I want to describe a few missing features around the transmission of open source projects to new maintainers.\u003c/p\u003e\n\u003ch3 id=\"discovering-active-forks-of-a-repository\" class=\"relative group\"\u003eDiscovering active forks of a repository \u003cspan class=\"absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100\"\u003e\u003ca class=\"group-hover:text-primary-300 dark:group-hover:text-neutral-700\" style=\"text-decoration-line: none !important;\" href=\"#discovering-active-forks-of-a-repository\" aria-label=\"Anchor\"\u003e#\u003c/a\u003e\u003c/span\u003e\u003c/h3\u003e\u003cp\u003eA lot of projects are maintained by a single person who, after a while, loses interest in the project, or even moves on completely and stops responding to any activity on the repository.\u003c/p\u003e","title":"Some missing features in software forges"},{"content":"","date":null,"permalink":"https://antonin.delpeuch.eu/categories/","section":"Categories","summary":"","title":"Categories"},{"content":"","date":null,"permalink":"https://antonin.delpeuch.eu/tags/","section":"Tags","summary":"","title":"Tags"}]