Markdown Version | Transcript | Session Recording | Session Materials
RASPRG
Summary
The Research Analysis Standardization Processes Research Group (RASPRG) met at IETF 126. The session focused heavily on the impact and measurement of Artificial Intelligence (AI) in SDOs, statistical analyses of Internet-Draft lifecycles and community mail archives (RIPE and IETF), and a proposed framework for analyzing data across SDOs. The chairs proposed holding an interim meeting to further explore the policy and technical implications of AI in standards participation, which received strong support from the room.
Key Discussion Points
Welcome and Introduction
- Slides: 0_rasprg-ietf126-chairs
- Ignacio Castro and Alvaro Retana opened the session, presenting the Note Well, IRTF policies, and the session agenda.
AI in Standards Participation
- Presenter: Mark Nottingham
- Slides: AI in Standards Participation
- Key Details:
- Mark outlined the double-edged sword of Large Language Models (LLMs) in standardisation. While AI aids translation, accessibility, summarizing, and editing, it presents existential risks: volumetric attacks (e.g., 60 AI-generated drafts in a row), wordy/autonomous mailing list bots, and lowered friction changing the balance of power.
- Tooling: Mark developed IETFLM, a tool written in Python (using Claude) acting as a local/cloud Model Context Protocol (MCP) server. It ingests mailing lists (stripping signatures/quotes), meeting materials, transcripts, drafts, and IESG ballots, providing embeddings for semantic search.
- Skills/Guidelines: Custom markdown instruction files for LLMs were created to guide AI behavior regarding IETF norms (
IETF interpreting,IETF contributing, andHTTPstyle guides based on BCP 56bis). - Policy Implications: The community needs to define limits on automated participation. Mark questioned if autonomous AI participation should be treated as disruptive conduct, and whether proof of humanity/raised participation bars are needed.
- Discussion:
- Andrew Campling noted that SDO rules are historically built on an assumption of good-faith engagement, which is no longer safe, requiring a root-and-branch rule review.
- Dirk Kutscher proposed running experiments in the IRTF to test AI policies and practices.
- Tommy Jensen expressed concern about the talent pipeline and suggested using these tools for upskilling/educational dialogues to help newcomers learn IETF norms.
- Brian Trammell agreed that friction can be a positive filter for learning norms.
- Nick Doty raised concerns that making these tools available might accelerate the existential threat by making it easier to pretend to participate, reducing shared context and "skin in the game".
- Jay Daley highlighted that the most dangerous issue is real humans using AI to comment on topics they do not understand. He shared a check tool he authored: mailing-list-ai-check.
- Jennings questioned model agreeableness and cautioned against assuming only newcomers use AI, as long-term participants also utilize it.
- Carolina Caeiro suggested focusing on guidelines and proof of humanity rather than discouraging newcomer AI use.
Measuring AI Authorship of IETF Drafts
- Presenter: Jaime Jimenez
- Slides: Measuring AI authorship of IETF drafts
- Key Details:
- Jaime presented a hobby project designed to detect AI usage in drafts using stylistics (hedging, common feeling words, word length, frequency distribution) combined with an LLM ensemble ("LLM as a judge").
- Technical drafts and translations by non-native speakers are highly prone to false positives, meaning 100% detection is impossible. To mitigate this, the tool flags items for "review" rather than asserting AI generation.
- Evaluating -00 drafts from Jan to Apr 2024 showed ~88% in the clear and ~12% with AI signals. AI usage clustered in working groups discussing high-level use cases/AI and was lowest in low-level wire formats.
- Discussion:
- Colin Perkins suggested setting different baselines by technology domain (e.g., routing vs. security) to account for differing default writing styles.
- A Representative from CEU and Dynatrace asked about majority voting in the ensemble and suggested using Wasserstein distance to compare distribution changes pre- and post-2022.
- Susan Hares inquired about whether AI draft rates correlate with RFC production slowdowns, which Jaime confirmed he has not yet analyzed.
- Priscilla Anunciacao observed that AI checkers often perform poorly on historical, pre-AI texts; Jaime noted that testing older RFCs yielded very low AI probability scores (under 6%).
Mortality of Internet Drafts
- Presenter: Marcello Santos
- Slides: 3_Marcello_Santos
- Key Details:
- Marcello analyzed ~31,000 usable drafts and ~5,000 RFCs from 2005 onward to understand why drafts "die" or become RFCs.
- Key Findings:
- Overall conversion rate of unique draft lineages to RFC is ~21% in the median.
- If a draft starts as an individual submission, it has a ~7% chance of WG/RG adoption. If adopted (or if it originates in a WG/RG), the conversion rate to RFC jumps to ~71%.
- Median time to publish an RFC is 2.5 years, though it has increased by roughly 1 year post-pandemic.
- Drafts presented at meetings have a 25% adoption rate. WG chairs or directors as authors see higher conversion.
- Geographically, Asia and Europe have surpassed North America in total submissions, but Asia shows a lower conversion rate (10%) before WG adoption. Latin America has a lower volume but a high conversion rate (~30%).
- Marcello created an interactive BI analysis tool at
drafts2rfc.com.br.
- Discussion:
- Susan Hares criticized the assumption that the IESG review and RFC publication process acts as a constant variable, pointing out that complex sub-states and loops within the IESG review cycle and WG last calls are not accounted for.
- Dirk Kutscher cautioned against treating IETF and IRTF drafts identically, as many IRTF drafts are successful contributions without seeking RFC status.
- Colin Perkins noted that reconstructing draft histories is significantly hampered by inconsistent metadata in the Datatracker.
- John Levine added that "dead" drafts are not failures if they serve to document why a specific idea was rejected.
Community Vitality through RIPE Mailing Lists
- Presenter: Ilke Ilhan
- Slides: Community Vitality through RIPE Mailing Lists
- Key Details:
- Ilke analyzed RIPE mailing list data to assess whether participation is declining and if engagement is concentrated among older generations.
- RIPE mailing list traffic showed a historical trajectory mirroring the IETF's: a steady rise, a peak in the 2010s, and a subsequent decline.
- Concentration: High concentration exists, with 19% of contributors sending 80% of emails in the 2020s. Out of ~1,600 accounts, eight accounts sent 20% of the total emails.
- Demographics: While old-timers from the 1990s dominate the highest-volume tier, there is a large, moderately active tier with many 2020s newcomers.
- Topics: Engagement is highly topical. The "Address Policy" list peaked during IPv4 run-out and has declined, while "Members Discuss" (charging schemes) and "RIPE Atlas" (technical metrics) dominate the 2020s. Address Policy was the most effective list at converting newcomers into multi-list contributors.
- Conclusion: Shifting traffic patterns reflect changing external challenges (e.g., post-IPv4 runout) rather than a direct decline in community vitality.
- Discussion:
- Peter Koch emphasized that raw email volume is misleading; for instance, the Atlas list traffic is dominated by highly repetitive, transactional "credit requests" rather than technical discourse.
- John Levine raised the difficulty of tracking participants who change email addresses and noted the generational shift toward platforms like Discord, Slack, and GitHub.
- Nick Doty asked how researchers can effectively track displacement of conversations to non-email media.
Analysing Internet Standards Data
- Presenter: Colin Perkins
- Slides: Analysing Internet Standards Data
- Key Details:
- Colin presented joint work outlining a framework for analyzing standards development organizations (SDOs) as sociotechnical systems.
- Observability Challenges: While technical artifacts (emails, drafts) are easy to extract, invisible factors like SDO culture, informal negotiations, individual/organizational influence, and agenda-setting cannot be easily measured.
- Data Extraction Pitfalls:
- Entity Resolution: Reconstructing people and affiliations is extremely difficult. Organizations have hundreds of variations (e.g., 280+ variations of "Huawei" in Datatracker), and consulting or funding relationships are opaque.
- Process and System Deviations: Historically, many drafts and processes have violated RFC 2026 rules, making systematic modeling tricky.
- Mail Archives: Legacy email formats are often malformed, breaking standard Python IMAP libraries.
- Ethics & Privacy: Research tracking individuals' "success rates" or errata frequencies must be handled carefully under GDPR and institutional review.
- Discussion:
- Brian Trammell agreed with the draft's findings and suggested formally raising data-archiving requirements to the IETF tools team via the IESG.
- Priscilla Anunciacao supported the draft and recommended creating a "checklist" or data-cleaning guide for researchers to address common Datatracker pitfalls (e.g., handling participants whose registered countries change during meetings).
- Jaime Jimenez suggested using
rsyncrather than the Datatracker API for large-scale data extraction.
Decisions and Action Items
Session Polls
A poll was conducted to gauge interest in holding an interim meeting on AI's impact on the IETF.
- Poll 1: Interested in an AI Interim?
- Yes: 21
- No: 0
- No opinion: 3
- Total: 61
Action Items
- Chairs to take the AI Interim proposal to the RASPRG mailing list to coordinate dates, agenda topics, and speakers for a virtual interim session.
- Chairs to initiate a mailing list thread regarding the potential adoption of the SDO data analysis draft presented by Colin Perkins.
Next Steps
- Community members are encouraged to contribute to Mark Nottingham's LLM skill files and Jaime Jimenez's draft analyzer.
- Further discussion on data cleaning methodologies and tracking cross-platform communication (e.g., GitHub, Slack) to occur on the RASPRG mailing list.