**Session Date/Time:** 23 Jul 2026 14:30 [00:00:42] **Ignacio Castro**: Hello, everybody, and welcome to rasprg. We have a very tight agenda, so we are gonna try to go quickly over the preliminaries. So first of all, Note Well, we have intellectual property matters, which are reflected in the slide that you can see here. The RTF follows the IETF intellectual property rights, IPR disclosure rules, and you can see the relevant information in the relevant RFCs. Note that we do audio and video recordings, and whatever you are saying here, will be recorded for as long as the Internet last. We also have a privacy and code of conduct, and you can see relevant information regarding this in the relevant RFCs that you can see in the slides. The goals of the IRTF are not about standardization, and we do not try to standardize anything here. Whatever we are discussing here are, is research related, that that might inform future discussions. About rasprg, rasprg is the research analysis standardization processes research group. We try to understand our to understand to understand the standardization process. And, the outputs are joint reports, papers, tools, data, open source software. It is not our goal to do hierarchical comparisons between SDOs or directly ITF operations. However, what we say here or what we what we do here, is helpful for ITF operations, that's very welcome. Alvaro and myself are the chairs, and you can find the charter in the data tracker. Please, use the QR code. Register. Use Meet echo for q and a. Speakers, please be ready. We have a little bit of a tight agenda. Is there any volunteer to take notes of this meeting? Thank you very much. That's very welcome. Could you please say your names for the record? Because otherwise, I will forget. Thank you. Right. And we have already saved a few minutes from the welcome agenda, which, hopefully, we will be able to use for more meaningful discussions. So we have a brief tight agenda with two different blocks. First, we are gonna be talking about AI in the standards process. First, say Mark Nottingham from Cloudflare, and then we say Jaime from Ericsson looking at two different aspects of using AI in standards. One, how to use it, another one, how to identify it. Then Marcello Santos from the Federal Institute of Sertao Pernambucano, if I say it correctly, is gonna be talking about the mortality of Internet drops, and then we are gonna compensate the mortality with the vitality of the community from Ilke Ilhan from Ripe, who is gonna be analyzing the Ripe community. And finally, we have, Colin Perkins from Glasgow University, with a proposed draft, on how to analyze Internet standards development organization with which for full disclosure, I'm also a coauthor. Thank you very much, and that's all from the chair side, Mark. So let me pull up your slides. [00:03:59] **Mark Nottingham**: Hello? Okay. I was hoping to do a hierarchical comparison with another SDL, but instead of that, I'll talk about AI, I guess. So I'm gonna talk about AI in standards participation. So this is in the process, not I I don't talk about drafts, creating drafts, or or or or AI in our protocols. It's more about using AI in the process itself. Next slide, please. [00:04:26] **Speaker 2**: Do I have a picture, I think? Yep. [00:04:29] **Mark Nottingham**: Lovely. Right. So I I don't think this is a surprise to anyone who's spent much time in the IETF recently, especially on the mailing lists. Increasingly, we're seeing people using AI as part of their standards participation process. They're doing it to write emails, to implement things, and to participate in the standards efforts. And in many ways, that's great. You know, I and and as I will show you, people are building tools and using AI to help them understand what's going on in working groups, to save time, to be more efficient, to broaden their their knowledge of what's going on in in the ITF. And, also, you know, you can look at AI as a as a very important equity and accessibility tool. People can use AI to translate if if English is not their native language because as we know, we use English for everything in the ITF, for example. And so these are are really great things that we see, and it's making our work broader and more efficient. However, we're already starting to see some problems. We're seeing groups that are complaining about being flooded with AI generated drafts. I heard one working group chair this week complaining that his group had seen 60 AI generated drafts all in a row, and that's very distracting for the group. We're seeing in some mailing list people sending really voluminous AI assisted messages where it blows out the size because as we know, AI can be quite lucacious. Yeah. I I actually asked for one of these slides, the previous version AI, to give me a list of of of synonyms for for wordy, and boy, did it deliver. Worse than that, we do see once in a while now either semi or fully autonomous bots on mailing lists, and it's pretty clear that somebody has said go and represent this position or push this draft in this working group. And you get these messages. You know, every single thread has a response of, yes. My draft can handle that blah blah blah blah blah. Extremely distracting, extremely disruptive to the work. And and that that's for a few reasons. Also, I down at bottom here, do mention that, you know, other folks have pointed out that having our understanding of proposals and our understanding of the discourse in in the ITF mediated by LLMs is also potentially problematic. I don't address that much in this talk, but I think that's something we should keep in mind. And and I I think that that you can see this even as as an existential threat in in that using LLM significantly lowers the barriers for participation. Now it is very easy to create human looking text. And while in the hands of someone who has good intentions and is using it to sharpen their arguments or to, you know, translate something to English or or something like that is is a powerful tool, in the hands of someone who just wants to throw spaghetti at the wall and see what sticks, it just gives them a volumetric attack. And especially for bad actors, this can be really problematic. We're seeing from and and, again, this is qualitative, but I see a lot of people who kind of have a shower idea. They say, oh, I've got a great idea. I want it to be considered by the ITF, and they use an LLM to sketch everything out and throw it at a working group. And then the working group has the task of digesting that and handling it with the due gravity that proposals that people bring here deserve. And and the problem is is that, you know, we have a a volunteer group of people who are protocol experts, who provide the, expertise and the experience, that create our output. ITF's legitimacy is not based upon its input. You know, we're not a representative organization. We don't have representative of countries or anything here. Our legitimacy is based on our process and our output, and and that means that we need to produce high quality specifications. But if the people who help produce those high quality specifications are inundated by a wave of AI slop, we have a problem as an organization. And and that's a big vulnerability for the organization, I think. If I put my corporate head on, which I do ever so rarely, I can't advise my company to bring work to a venue that's a circus. And and and I feel like that's the direction we're headed in if we don't do anything. So and here's a quote. I was talking to John Peterson about this earlier this week, and I think John sums it nicely. You know, rough consensus and running code are predicated on us having humans coming to that consensus and humans running that code. And AI changes the friction of both of those acts. It's not that it's always bad to use AI, but when it's low friction, it changes the system. The checks and balances for making sure that we have good outcomes change. And so this is a threat to the balance of power at the IETF. So when I think about this, I think about what carrots can we provide? What can we you know, how can we incent good behavior in a positive fashion? I will get to a different emoji later in the talk. So first of all, I've been working on a tool called IETFLM now for a couple of months. It's been through a number of iterations. I'm not going to give a demo. I'm sorry. Bad experiences. So, basically, ITF LLM allows you to gather a corpus of of materials from the ITF. It's about 22,000 lines of Python that that Claude very generously wrote for me. He's very kind. It can run either local on your own machine or or it can run-in the cloud. You can host it on Cloudflare. You can host it on Akamai. Can host it on AWS. It just needs an instance and some like, s three or something s three ish as well as the ability to do embeddings. It has three modes of operation. You can yeah. Like I said, you can run it on your own machine as an MCP server. You can run it as a cloud MCP server, and you can also this is the older mode of operation. You can gather a a folder of files and then point a tool like NotebookLM, if you're familiar with that, at them, and it does the magic. I'm probably gonna deprecate that last one because I think the MCP server is is a lot more powerful now than than that. And it gathers meeting materials. It gathers the the mailing lists using IMAP. It creates an syncs to a local cache. Transcripts from Ecker's wonderful transcripts that he's creating for us. The issues list, if they're available, it figures out where the issues list are based upon the data tracker page. The drafts that are being discussed, IESG ballots, the relevant RFCs, all of this is munched together, and it does a tremendous amount of processing on it to make it much friendlier to the LLM. So instead of just giving you, you know, raw email, it strips quotes, it strips signatures, it creates markdown files, and it creates the relationships between the emails and creates threads. So the the LLM can natively digest the threads, for example, for email. And then it goes and creates embeddings across all these things. So the LLM can do semantic search using the MCP server on that corpus of information. And what you find when you're using it is that, you know, all this information is available to your LLM of choice, Claude, natively. It does the web search tool, and it knows where the ITF is. It it it does know about us. And and so you'll say, you know, tell me what's going on in this working group, and it'll go off and do a couple of web searches and give you a canned result. What you find when you use this tool is that the the answer you get back is much, much richer because, effectively, the LLM has a budget. You know? When it's doing web search, it it says, okay. I'm gonna do three or four searches, and then I've used too many network calls. I've used too much time. I've used too many tokens. I'll give you an answer based upon those three or four searches. But when it has a an MCP server that gives it a very high efficiency interface that's tailored through a lot of iteration to the use cases we have for this tool, you get a lot more complete and accurate answers. And it's really nice to see it when it's working well, which usually it does. That's not why I'm not giving a demo. So what can you do with it? Here are some examples that I've been using it for. You know, we the first one, we actually in AI Pref had some fairly contentious discussions, and someone using the tool went off and said, I wanna make a proposal for this particular issue, and I want to consider all the arguments that have been made to date. Can you come up with a proposal? And they had a long chat with I forget which LM they were using. And they came up with a a reasonable answer that actually got adopted by the working group as the basis for its its solution for that problem, which I think is great. And that was something that hadn't occurred to a human. You can also just kinda keep up with the working group and say what's being discussed currently in this working group or what happened in the last three weeks. When you get these voluminous email lists, that's actually really useful to kinda get those summaries. You know, given the objections to proposal, what, you know, what what can I anticipate from my proposal? The the second to last is really interesting. I did this a while back. I asked it to review, our chairing of the I preferences working group and got a report back of how biased we were as chairs. I think it's a soup super useful tool for self evaluation by chairs, and it it luckily thought we were pretty good. [00:13:41] **Marcelo Santos**: So [00:13:44] **Mark Nottingham**: and and, of course, of course, the first query that everybody these days is gonna ask, tell me what's going on the TLS working group re on a recent. If you don't know, look it up. So second of all, that's just an MCP server that gives you this ability to to access this this rich corpus of of material for working groups and and research groups. And by the way, it is not just IETF. It's also IRTF. You can just tell it go keep up with the last call mailing list for me if you want or the IETF list, although you're welcome to subject yourself to that fire hose. What what also became necessary was some some context about the ITF itself to give to the LLM to guide its its thinking or sorry, thinking about the ITF. And so I created some skills. If you're familiar with skill files, it's just some markdown to tell the LLM how to do stuff, basically. So the ITF skill is assistance for participating in the ITF. It's right now three skills. It's all marked down. It's not any scripts or any any helpers yet, and this is how you install it. Don't worry. I'll have this at the end. All of these things there in these skills are just my personal opinions. Based upon my experience in the ITF, I'm happy to take comments and issues and pull requests. The first one is IETF interpreting. So it's basically norms for reading what a group decided before you claim it. So, you know, this this is based on all the different interactions I had with Claude and and how it made assumptions about the IETF that weren't really warranted. So saying, you know, we don't vote. And just because it feels resolved on the list doesn't mean it's actually been called for consensus and that we're all out participating as individuals, but there's some nuance to that, things like that. Teaching about the RFC streams, the fact that if you've Internet draft, it has no actual status in our process unless it's been adopted by a working group. Things like that. The second one, and which is a little more interesting, is ITF contributing. If if you're an LLM and you're contributing on behalf of a human to a working group, especially in the email, what should you be doing? And so it's really trying to drive home the accountability aspect of that that, you know, don't operate in an automated fashion. You need to make sure the human understands what it's sending and then, you know, work through that with the human. Ideally, disclose AI involvement. I think that's a great subject for discussion because there's pros and cons to that, especially the register of the communication. As we know, LMs are quite wordy, and so, you know, telling it to to respect people's time and attention and just stick to the facts, give examples where it's necessary, but don't go and do the thing that LLMs do. I I've processed a number of messages that really looked AI generated from the lists through this through these guidelines, and it helps a lot. It cuts down paragraphs to a sentence or two, and that's what we, I think, we wanna see. Granting claims in in real things, the nature of consensus, stuff like that. And then finally, the third skill is, specific to HTTP. We have an RFC. BCP 56bis is using, HTTP to build new protocols, building protocols with HTTP, and this is kind of the guidance of how to build, you know, restful protocols in the ITF. We also have an HTTP editorial style that we apply to the HTTP working groups drafts. And I I chair the HTTP directorate, which is the review directorate for HTTP specifications. And it's a difficult directorate because often a group will come to us at last call and say, hey. Look. We did something with HTTP, and we'll say, that's a shame because you're gonna have to rip it down to do it right. So we wanna give people tools to make this easier. It's very early days, but it's very promising. I've already had a few interactions with working groups using this skill, and it's amazing. You say go review this draft. It looks through all the advice, and it comes up with a list of of proposals and and recommendations and concerns about that draft. It's not a 100% correct, but it gets you maybe 70 or 80% of the way there. And in the right hands, it gives the reviewers a a really powerful tool to take that initial burden off of their review. As the person who assigns the reviews, I always die a little bit inside because I feel like I'm, you know, screwing somebody's week by giving them this voluminous document to read and and then trying to argue with someone about. Yeah. I think it also can help groups. I think, eventually, we wanna get to a place where we give the skill to working groups and say, you run the skill a couple of times. And when you have questions, when you have things that are ambiguous, come to us, and then we'll help you, you know, with humans and the experts. Yeah. So the other emoji. This is the closest I could come to stick. I think we need to discuss policies for AI involved AI assisted or or automated participation in the IETF. The tools are probably not enough. The carrots are nice and everything, but, you know, there's lots of questions. You know? Is pointing an LLM at a list and asking it to argue a point for you autonomously disruptive conduct? I would say it probably is, personally. I we don't have that documented anywhere. We don't have any norm for that or any rules for that. Is stating that you believe another participant is an auto autonomous LLM, so you're gonna ignore their input. Is that harassment, or is that unprofessional? Some people might think so. I don't know. That is one possible way to check that kind of behavior, though, so I think we should think about it carefully. What if it's someone who's just being really sloppy and they're used to the LLM, but it's still kinda human guided? Is that okay? I don't know. I've heard people talking about disclosure requirements that you require people to disclose their use of AI. I'm not sure that's gonna be productive because people have a strong incentive not to disclose, and they're gonna try and get away with it. Maybe if that becomes normal, if if everybody's doing it, it might be okay, but I think we need to have a discussion there. I've heard people suggest maybe we should consider how we consider new change how we consider new work. So instead of allowing people just to post a zero zero draft and forcing the working group to look at refuse any examination of new drafts until we actually agree on use cases and requirements and have that discussion first. Because we do see a lot of people just throwing ideas over the wall to see if they stick. We've I've heard people talking about openness of the community. You know, a lot of these problems are coming about because it's really low cost to join the IETF. And despite Jay's best efforts to get a data tracker account and then throw stuff at us, do we need some sort of raised bar for participation in the community? I think that it's gonna make a lot of people deeply uncomfortable and change the nature of the community, but that's one of the defenses I've heard people talking about. And then to me, I like you know, there are the rules, which a lot of these enforcement is gonna be a problem. I think norms should come in this into this discussion as well if if, you know, everybody is disclosing a certain way or everyone's behaving in a certain way. And and if especially, it's very clear what the expectations are in terms of use of AI that hopefully will encourage good behavior, at least in the common case. I'm not sure it's gonna be completely successful because in my observation, a lot of the people who are coming in and using AI are newcomers to the ITF. They're like, oh, I used to not I used to be scared of or didn't have the energy to make a proposal to this group, but now the barrier to entry is drastically lower, so why not? And they may not care that their reputation in this community might be trashed if they use AI incorrectly. To them, it's just throwing it at the wall and seeing if it sticks. So we need to think through those cases as well. Yep. These are the links. They're on GitHub. They're in PyPy. Pretty easy to install. The only thing I will say here is these projects are moving extremely fast. So if you do play with them, update very, very often because I'm I'm changing them pretty much every week, if not every day. That's all I got. [00:21:22] **Ignacio Castro**: Thank you very much. Questions? [00:21:25] **Andrew Campling**: Yes. Andrew. [00:21:28] **John Levine**: Oh, good. [00:21:30] **Andrew Campling**: Hi. Andrew Campling. Really interesting talk. Wench and Dirish, I wasn't expecting, but that's that's good. On the on the sort of stick thing, I think we're sharing a problem or experiencing a problem that the RARs have had relatively recently, but from a different completely different angle, which is I think our rules are crafted on the assumption that, basically, everyone's engaging in good faith for a particular definition of good faith. And what this is really a symptom of is we we need to reconsider the rules from the ground up because we might reasonably conclude that that's not a safe assumption anymore, not specifically just because of AI, but including AI. Which so, basically, to sort of pun your your your image. We, yeah, we need a root and branch review of all the rules. Because, as I say, just like the RARs, they can't it's not safe to assume good faith engagement. Therefore, we have to do things differently, and that's gonna be hard. Yep. [00:22:39] **Mark Nottingham**: I tend to agree. I think, yeah, more robust processes, and you need to look at the political economy of it. You know? It's it's really there's a this is a really fundamental change. [00:22:49] **Dirk Kutscher**: Dirk? Hi, Dirk Kutscher. Thanks, Mark. This is really great. Just on the policy considerations. So what we in the IRTF can also do is to run experiments. So, normally, we do this for technical specifications, But we could indeed also have a group that experiments a little bit with practices and maybe learn some potential useful policies that could be recommended. Okay. [00:23:14] **Mark Nottingham**: I'll I'll propose an experiment. To participate in the in the research task force, You need to give me a $100 to come in the room. I think that'll solve the problem. [00:23:23] **Ilke Ilhan**: Right. [00:23:25] **Mark Nottingham**: There there's a lot of experience behind that technique. [00:23:28] **Dirk Kutscher**: You know, I'm I'm I'm looking for successful experiments. Oh, define [00:23:33] **Speaker 2**: success. [00:23:36] **Dirk Kutscher**: And on the general tooling question, so I think, I mean, this is really good, really good start. And I think this could, of course, should be taken further. And I think there could be, and you are familiar with that, like, you know, different, let's say, ways of using these tools. So for example, you know, in like, inside facing usage for our own community, but maybe also interfaces to the public. So how can they better understand, like, what we are doing? And I think it's it's yeah. It could be quite quite fruitful to just, you know, experiment with these ideas. [00:24:10] **Mark Nottingham**: Right. I I think the interesting thing is, you know, a tool like this, the the LLM, could make it a lot easier for policymakers to understand what [00:24:18] **Dirk Kutscher**: we're doing. [00:24:18] **Mark Nottingham**: That that's very interesting. Yeah. [00:24:20] **Dirk Kutscher**: Thank you. [00:24:23] **Tommy Jensen**: Tommy Jensen in. I was sent here by commentary from the plenary when talking about the responsible use of AI in the IETF. Really appreciate the talk, but one of the things that popped in my mind was a net new fear about the use of AI, which is it would be great to be able to upskill people to participate with us. What do we do about the already challenging talent pipeline of getting people interested in our technologies and our leadership positions If they can come in, accomplish a goal, and leave having just used some tools, to that end, have you considered or looked at having the interactivity of these tools be a teaching opportunity for the user? [00:25:05] **Mark Nottingham**: I've started to do that in terms of the dialogue that it encourages the, you know, the it tries to actively discourage the just one and done kind of, you know, throw this at the working group interaction. It's more it it encourages, the LLM to have a dialogue with the user to talk about the requirements, talk about the existing work in the group, and engage in that fashion. So maybe we can go further in that direction. [00:25:25] **Tommy Jensen**: Yeah. That would be cool because if there was a, shared concept of terminology and framing, like, starting to teach everyone or get us on the same page on what a use case is versus what a a goal is, that'd [00:25:38] **Mark Nottingham**: be lovely. I think we probably need to be more explicit in the, in the ITF about that methodology as well. We probably need to do something like write down what we mean by consensus. [00:25:48] **Brian Trammell**: Yeah. Hello. Paging Pete Resnick. Hi. Brian Trammell. Plus a lot to a lot of what Tommy said. I think the way I would encourage us to think is tastier carrots and bigger sticks. Like, I think one of the upsides of the friction of of, participating in this community is you actually do get to learn the norms and you understand the norms before you start to contribute, and that also leads to people sticking around for longer that we don't scare off. I see a lot of people at this meeting who I've never seen before that makes me very excited. The barrier to entry is lower. Is there a way we can take the tools to take that low barrier of entry and turn that into a feedback loop? I have no idea how to do that. I'm willing to try and let you help. I'm playing with the skulls right now. [00:26:35] **Mark Nottingham**: Thank you. Good. I like how you describe it in terms of friction because that's how I think about this a lot. And, you know, we wanna reduce friction for the things that we wanna encourage and make sure that it isn't accidentally reduced for the things that actually make our systems work. [00:26:48] **Jaime Jimenez**: Yep. Yep. [00:26:52] **Alvaro Retana**: Nick. Go ahead. [00:26:54] **Speaker 2**: Yeah. Maybe maybe to follow-up on that. I I I certainly agree with your diagnosis of the problem, Mark. I I agree it's existential for, I think, many open standard setting organizations. I I was confused then by the presentation of the tool. It seems like you're putting a lot of work into into tools that will accelerate that that particular problem. That, like, hey. We we we can make it even easier for people to attack the IETF and W three C and others if if we give these tools. I I don't know if you're just trying to, like, make sure we confront this problem right away by by making it as easy as possible, but [00:27:39] **Karuna**: yeah. [00:27:41] **Speaker 2**: I I I guess I'm I'm just struggling to see how providing the tools is just not making it easier for people to pretend they're participating without participating. And and I think actually that gets to some of the other things is that I think you're saying, well, there's some good uses and some bad uses. I'm not sure I agree with you on the good uses. Hey. You you can you can, like, analyze arguments, with without talking to someone in the mailing list, or you can follow a a group without, showing up at any of the meetings or reading any of the writing. Those sound like an efficiency goals, but but I actually think, like, they they decrease effective participation. If if I don't, like, ask anyone, hey. Is this a good idea or a bad idea or why? Like, we we lose the sort of shared state of, like, talking through ideas and understanding why they're good or bad. Or if I don't have to show up at any meetings, or engage in any conversation or read what anyone else is writing in order to to know what's going on, well, then there's there's, like, less less community commitment. There's there's there's less skin in the game. There's less [00:28:49] **Ignacio Castro**: Sorry, Nick. We are running a little bit tight on time, and I think that the point is clear. What are the negative sides and that that you feel that this might help those negative sides? So maybe it could be good if we let Mark answer because there are a number of questions after. Sorry about that. [00:29:04] **Mark Nottingham**: No worries. Nick, I yeah. I I hear you, and I've I've I've thought these thoughts as well. What kind of drove you? I'm I'm not a big AI booster, but it it became violently clear to me than the working groups that I've been participating in that people are already doing this a lot. And I came to the realization that if they're gonna be doing it and we can't stop them, we might as well give them tools that guide them on better paths and and hopefully guide them towards the discourse that we need to actually come to consensus. And and that's really my goal with this is to give them good quality information and good guidance so they can engage in the process with these tools in in a productive way. Now if we wanna ban AI altogether, great. You know, I I I'm interested to see how that's enforced. But if not, we need something to to help people. [00:29:53] **Jay Daley**: Jay Daley, fellow AI enthusiast. [00:29:56] **Mark Nottingham**: So Hold on. [00:29:58] **Jay Daley**: Yeah. A few slides ago, you were enumerating behaviors, improper behaviors probably about the use of AI. Are you planning to turn that into anything to, you know, sort of a list of these are the things we know that people are doing that you probably shouldn't be doing? [00:30:17] **Mark Nottingham**: I I believe the IASG is interested in this topic, and I'm happy to have a chat with them. So I think it's probably gonna come up tomorrow. [00:30:23] **Jay Daley**: Fantastic. So the one I would add then is what I think I'm seeing, and this is all suspicion, is people participating, real humans using AI to participate with on a subject they understand absolutely nothing about, not even close. You know? And yeah. Yeah. So which I think is actually a bigger issue than many of the others are saying. [00:30:44] **Mark Nottingham**: I tend to agree because it it you know, the it's asymmetric, and that's really a lot of the base of the basis of the issue here. [00:30:49] **Jay Daley**: Yeah. Yeah. Thanks. [00:30:52] **Jaime Jimenez**: Hi, Mark. So I took notes just to remember. So on the efficiency tool so, I mean, is it an efficiency tool at heavy users, in my opinion, quickly gets saturated with the actual draft writing processes? So you may see more impressions to your SEO drafts. In my opinion, the impact of AI is that the consensus process that happens here no longer really needs to happen here anymore. That that could be more concerning, but that's maybe for a side discussion. Then on the tooling, these tools, like the one you just showed, like, most r and d teams already have this type of thing already or they have built it by themselves for a while. So there's no stopping this, I think, in my opinion. Then on the next presentation, I will show, like I mean, filtering version zero zero drafts could actually help a lot if you have automated tools that could flag those, that could actually decrease the burden for for reviewers. And then the last one, I'm sorry if some of my my emails have sounded a bit AI like. I I use it a lot. I don't know. I apologize before. I I don't know if it's me, but just in case. [00:31:52] **Jennings**: Thanks. Sorry, Jennings. Department for business innovation science and trade. I have two questions and one comment. So the question is, do you have a sense of how agreeable the model is in terms of its response to queries? Because I can imagine one of the issues that could come up is it being generally positive and receptive to ideas and then people coming to the community and finding that they are perhaps slightly less agreeable than is otherwise perceived. The second is how far back does the corpus go? And is this something that is kind of useful for a institutional memory storage of some capacity, particularly thinking of people kind of aging out of the IGF and are not returning and also losing a little bit of that? And the third thing was on the, I guess, a comment on particularly the response to people using AI and LLMs in their contributions and whether that should be considered kind of harassment or a code of conduct violation. I do think we need to consider that norm quite carefully, particularly in the context that actually the use of LLMs is not exclusive to new participants, and it would be a mistake for us to consider that that is the case. There are lots of active long term participants in the I two f who have established reputations, who are using AI to wholesale produce drafts, who are not getting leveled with the same criticisms and critiques whose ideas are being taken seriously. And this becomes a risk of creating additional barrier to entry. [00:33:07] **Mark Nottingham**: Yeah. Yep. On that last one, it's, yeah, it's it's how people use it, not whether they use it. Yeah. And I'm sorry. Should I answer the questions? Or I don't know. Because I already forgot them. Sorry. I it try and beat it out of it. I'm sure I need a bigger stick. I found it it especially with with characterizing what's already happened, it it usually sticks to the facts. I mean, of course, you have to check everything. This is the nature of using LM. But but, yeah, we can definitely make the skills more you know? How far back it goes, you gather. You you direct how far back it goes, basically. I think it has the capability to go back because of the way that the data tracker has kept information so forth, roughly 2,000 or so. So it's it's pretty [00:33:52] **Colin Perkins**: good. Yeah. [00:33:54] **Jaime Jimenez**: Yeah. Yeah. That. [00:33:55] **Colin Perkins**: Hi. Calling Perkins. Yeah. The data tracker is useful back to about 2005 or so. It gets a bit hazy earlier than that. Fully agree with all those comments, particular for new participants, especially those for whom English isn't a native language. We need to be very careful because this is probably a very useful tool for some of them. We have an awful lot of guidelines RFCs if we have. A lot of them even provide useful guidance. It might be an interesting experiment to maybe for this group, maybe for another to try and put those into a useful form for the airlines to ingest and see if that is helpful to the review teams and to office. [00:34:37] **Mark Nottingham**: I would love help with the skills because, like I said, they're just my opinions right now. I'm sure they can be expanded. I'm sure they can be based upon more solid materials. And in my ex like I said, the the BCP fifty six one, it does a pretty good job. You put the stuff in there. You you you organize it in a way that's more amenable to the LLM, and it it does follow the advice. So, yeah, let's let's get it done. [00:34:58] **Colin Perkins**: It's possibly something we can leverage the, direct to help. [00:35:02] **Speaker 2**: Yeah. [00:35:06] **Ignacio Castro**: Karuna, if you can be very quick. I I understand everybody wants to discuss. We might need to move this to an interim at some point. [00:35:11] **Karuna**: I wasn't sure if, I think I was beyond the lock, so I appreciate the [00:35:15] **Ilke Ilhan**: the [00:35:15] **Karuna**: opportunity. Three points very quickly. I was interested, excited to learn the ISG is doing something on this. And I'd be curious to hear more from Tommy actually on sort of what it is that you're doing and how we can engage the community since we're having these conversations. On the newcomers and them using these tools, I actually think we don't need to sort of frown upon the use of AI by newcomers. And it gives me sort of a mix of, like, you know, oh, the next generation is doomed. They're not gonna work as hard as we are. And also this idea about, like, the technology is gonna sort of spoil them. Like, you know, video games will bring ruin the sort of the kids, and I think it's it's not the case. I think it's all about guidelines. And I think that's where the ISG can really sort of signal to the community, like, the expected use of it. And then I just wanted to say in terms of, like, who participates in ITF, I think we should move to proof of humanity and not have bots riding on many lists. [00:36:25] **Mark Nottingham**: Just to the newcomer point, I think the issue is not so much, you know, about the newcomers themselves. It's about institutional capacity. You know, we we have the capacity to process a certain number of standards, to have a certain number of documents going through the RFC editor queue, have a certain number of working groups here, and have a certain number of views at the same time. And when you lower the barrier so much that, you know, the input pipe, the the top of the filter expands to you know, by several orders of magnitude. Even if a lot of those are good proposals, we don't have the institutional capacity to even handle that and and triage it. And and that that's a worry for me. [00:37:04] **Ignacio Castro**: Thank you very much. Thank you very much, Mark. And continuing with the topic, we have Jaime Jimenez on measuring AI alpha SIP on of ATF drops. Thank [00:37:15] **Jaime Jimenez**: you. So I'm not such a great speaker as Mark. I'll do my best. So, basically, I've been working in the past couple of months on a tool to detect AI usage on ITF drafts. It's running on this website. It's also an API access if you wanna do it programmatically. I do basically, this is the methodology. So I have a set of stylistic samples that I take. So I take chunks from each of the drafts. I cannot process the whole draft. In many cases, they are, like, 40 or 50 pages. It's pretty cost inefficient. So I just take a sub subset sample. I remove all the templated information, obviously. And then I have some so based on statistical statistical techniques to see the, you know, the type of wording you use. If you use longer words, smaller words, word combinations, you can plot them, then that could hint to to AI usage. Then that, I pass it together with the draft to a model ensemble. So it takes this concept of LLM as a judge. So you can actually ask your LLM for opinions and you pass hard data. The output tends to be more reliable. Ultimately, it still is an opinion from an LLM, you can none none of the techniques to my knowledge I read a few papers on this. And this was a hobby project, by the way. So, I mean, don't expect super deep academic purpose here. But to my knowledge, it's not really possible to to guarantee 100% that something has been AI made. Especially if you start curating it, it's really, really tough. In any case, for for basic stuff, for for things that have been a bit sloppy, then that that worked well. Anyways, so then there is a consensus signal, and that is what you see on the website, basically. So the algorithms are are borrowed from the telemetry and AI detection literature. They are at the at the end of the presentation if you are curious. As I mentioned before, each signal by itself is alone. When you put them together, especially with the stronger signal like hedging or or I don't I know the other term, but, like, with those, actually, it becomes a bit more precise. And then the the the drawback I mean, one of the problems we have in in the ITF in in a sense is that the document documents are technical by nature. It's not like a novel or or like a literary text. So it's much easier to get flagged as AI. And then for non native speakers, I mean, you you would, by default, get more chances to be flagged. And if you use a translation tool, for sure, you're gonna get flagged if you don't touch it. So that is also to be taken into account, which doesn't mean that you don't use AI. Just to be fair there. These are the style statistics. I mean, maybe I'll because we don't have so time, I will not go through all of them. But for instance, the first two hedging, common feeling words, like, you read load bearing or things like that, having a dictionary process in, for instance, in ID needs, I will mention it. That is trivial. We should add it in my opinion. The mean work length, as as Mark was mentioning, like, wording usually indicates AI. There's a frequency distribution. If you take all the words, you put them together in a graph, the slope is different for non AI drafts. I didn't do the work that I probably would do now of taking all of the drafts, version zero zero drafts before 2022 and getting a statistic statistical measures or or measurements from those. But, anyway, future work. Anyway, there's more on this. You get this type of nice report that you can read. If you want, you can also read the actual opinion of the LLM based based on the report. You may get mixed outcomes in all of them, and may they they may disagree as well. So this is also how you tweak it. You could you could for instance, Opus tends to be a bit more rigorous, so you could give more weight to Opus versus other models. You could just do a statistical measures. You could do other things. Then this doesn't work. Zero. No. So the the well, yeah, there there there is certain dimensions that I so you can you can ask a model that to return JSON, and then you can flag some you you can request certain data structures that then you can process. You can calibrate also. For instance, in my case, I I make it quite lenient, I'm not trying to to to trust versions here submissions, of course. And then when there is disagreement, I just mark it as review instead of as a as a flag. And they they, basically, they don't have a lot of argument arguments, statistical arguments. They tend to flag it as clear. Then you can I mean, in the website, you can see looks better on the laptop, but you can see the plot? I I did only from January to April. I did notice a slight trend of AI usage right before the meeting, for example. But overall, it's kind of balanced. From what Mark said, maybe there has been a bit of a of a increase here. I don't know. Again, in retrospect, probably, would just run it for a much longer period and see what it shows. Maybe this is something interesting for IRTF, by the way, if you guys have budget on these things. If you wanna run it, we could do a a bigger experiment. The style deviation, so these shows, for instance, the these will be the drafts. The drafts that are in the clear, these will be flagged, and these will be clearly AI. So about 88% of the drafts are normal, and the rest are kind of show signs of AI usage to better or worse degree. In my experience, I use AI for almost everything these days, and I have I I know everybody's using it for many things. If you work with the AI, it's not difficult to reach the green pass, especially if you're doing technical documents and you are, you know, like, oh, you have a new C verse realization format and you know what you want, the user want the AI to generate quickly and you check it and so on, that will not be flagged. So these are probably, you know, not not the best drafts out there. And and yeah. So they have they have a strong signal for AI, but it could be that they are just translations. I don't know. I didn't went through all of them. That's another thing. If somebody has a team of volunteers that wants to verify, you know, human verification that helps also to train and improve the models and the and the the system. Then the by working group, this was pretty cool to to see. So you can also plot based on the draft submission name, which working groups or which names are they using. So, for example, you have working groups that are more about AI and about the use cases and more, like, higher level topics. They they are more yeah. They they seem to have more AI in them. And then much like these groups that are pretty specific wire formats and and byte layouts and so on, very concrete, tend to have far less AI. Then deployment wise, I mean, as I mentioned before, you have it in a in a website and you want to use the API endpoint. You just use the API. You can filter with jQuery and you just get the little bit that you want. Yep. That's basically how it works. Cost, not expensive. So the model ensemble I use is not not the not I mean, it's normal price. About I mean, for $35, you can do the 1,000. Of course, if you process the whole document, maybe it's three times that amount. If you use more expensive models, multiply by two and so on. But you could do pretty cool experiments with very low budget. Yeah. Yeah. And if you can use reasoning models, they tend to give also more verbose output. I haven't measured if this is better or not, to be honest. Wish list. So I think it could be better. So I think we could we could actually analyze every version zero zero by default. And in some cases, for instance, for the stylistic aspect is for free, and you could just give a warning, for example, where the user is submitting something that clearly has a hedging or common AI words. Then the yeah. I could also increase the the draft content that they pass on the on the model. I think, as I mentioned, like, the the website is pretty lenient. It's not trying to destroy people's reputation, as Mark said. But it could be also more strict, to be honest. So if we want to to give feedback to the to the author, like, hey. You really need to look at this. It could be also phrased in different ways. It could be showing what parts should be checked and so on. And then also the last the last bit. So English as a second language will have similar surface surface signals. Yeah. And that's it. Thank you. Thank you very much. Comments, questions? [00:45:56] **Ignacio Castro**: Colin, you're first. [00:46:00] **Colin Perkins**: Hi. Colin Perkins. Yeah. This is really interesting. If I understand correctly, you're putting all you're putting a bunch of drafts from across the ITF into this. [00:46:11] **Jaime Jimenez**: That's I couldn't [00:46:12] **Colin Perkins**: If I understand correctly, what you're putting into this is a bunch of drafts from across the ITF to get a a single statistical model of [00:46:21] **Jaime Jimenez**: All all yeah. It's analyzing, stylistically, every single version series you draft since January. [00:46:26] **Colin Perkins**: Yep. What I think would be really interesting would be to look to see if there are differences across the ITF in the typical writing style. Are routing drafts written in a really different style to security drafts, for example. [00:46:40] **Jaime Jimenez**: Yeah. Like, the the website [00:46:42] **Colin Perkins**: in terms of AI use. [00:46:44] **Mark Nottingham**: Right. In [00:46:44] **Colin Perkins**: terms of the baseline. [00:46:46] **Jaime Jimenez**: Right. I see. Different baselines for different domains. That could that could that could be an interesting approach. Yes. I agree. [00:46:59] **Representative from CEU and Dynatrace**: Yeah. I'm Central European University and Dynatrace. Thank you for the presentation. It has been very insightful. I have a couple of questions, and then you can try to answer all of them or we can also discuss later. So regarding the ensemble model use, I wonder whether you use majority voting or something else because in my experience, some models perform better on certain type of, let's say, sentences and some on some other type of sentence, for example, like short versus longer ones. [00:47:27] **Colin Perkins**: Mhmm. [00:47:27] **Representative from CEU and Dynatrace**: Then the second point regarding statistical evaluation metrics, did you try to use Wasserstein distance to see [00:47:35] **Jaime Jimenez**: the cost, yeah, of bringing [00:47:37] **Representative from CEU and Dynatrace**: the the two distributions together. And maybe my third question, feel free to let me know if I misunderstood. But based on the picture you show the graph you showed, it looked as if difference in differences, I think it was the previous one, might could be used on that. And I was wondering if it would be possible to take the IT of drafts before the AI was Yeah. [00:47:59] **Colin Perkins**: That's what [00:48:00] **Jaime Jimenez**: I mentioned. [00:48:00] **Representative from CEU and Dynatrace**: Synthetic data. Ah, you mentioned that. [00:48:02] **Jaime Jimenez**: I meant before 2022, which I assume that's when AI started. [00:48:06] **Representative from CEU and Dynatrace**: Okay. But would it okay. Thanks for that. But would it be possible to create synthetic drafts based on the ones before the AI and then have parallel trends in place [00:48:19] **Jaime Jimenez**: and compare. [00:48:20] **Colin Perkins**: I don't [00:48:20] **Jaime Jimenez**: know if synthetic I mean, so to answer the question, so the first one on the models, I I'm not very sophisticated. I use I mean, some of them are of of the same branch, you know, like Sonnet and Opus and so on. So I know Opus is more rigorous, it seems. The second I forgot the other. Sorry. The second [00:48:37] **Representative from CEU and Dynatrace**: The last thing? [00:48:37] **Jaime Jimenez**: Yeah. That one. Yes. And then the last one, I forgot as well. Sorry. [00:48:41] **Representative from CEU and Dynatrace**: The last one so taking [00:48:43] **Jaime Jimenez**: I'm terrible. Yeah. Yeah. The the yeah. Yeah. So one one of the issues I so now in retrospect inside, I realized that I should have probably done some statistical sampling of the drafts, version series of drafts before 2022, maybe by working group or by domain, create some sort of metrics for that, typical words or typical dis like, still metrics data, and then compare with the current one. And probably, you'll be the results will be probably worse, I would guess. And then on the so the results, even though this is red and it's likely AI, I don't have a way to truly, truly say it is AI because I don't know. So that is the all the other problem. So maybe human based feedback or synthetic data to generate drafts that I mean, I don't know. There there could be better ways to do it. Yeah. That certainly. [00:49:33] **Ilke Ilhan**: Thank you so much. [00:49:34] **Jaime Jimenez**: Thank you. Jay. [00:49:36] **Jay Daley**: Thank you. Jay Daley, so three questions for you. Have you compared your statistical analysis against any of the commercial models such as PanGram or GPT zero? [00:49:47] **Jaime Jimenez**: Well, yeah, mean, I use g p t GPT o? [00:49:50] **Jay Daley**: No. No. No. No. So there are there are public services that are AI detector services No. Or API. And they they do similar things, but they charge No. [00:49:58] **Jaime Jimenez**: I I if Okay. If there is interest, I am more than happy. [00:50:01] **Jay Daley**: No. No. No. Just curious. Do you do any preprocessing of the IDs to remove certain signals before doing [00:50:09] **Colin Perkins**: all the [00:50:09] **Jaime Jimenez**: I I remove only the boilerplates. [00:50:12] **Jay Daley**: Okay. So in my usage of those commercial services, there are certain things that need to be removed in order to get a good signal from them. Okay? Signature blocks, various other things confuse it, salutations, and things like that. Odd odd things. [00:50:26] **Jaime Jimenez**: I mean, drafts usually don't have Yeah. No. They're very they know. [00:50:29] **Jay Daley**: Salutations, but the the the author's box and stuff really [00:50:32] **Jaime Jimenez**: Yeah. All the template did is removed. [00:50:34] **Jay Daley**: Alright. It's great. Okay. And then On the boilerplate. Sorry. Yeah. That's right. And then my third question jumps [00:50:43] **Marcelo Santos**: out of my head now. Okay. Great. [00:50:44] **Jaime Jimenez**: Thank you. We can we can talk after. It's okay as well. So that's it? Ano? Sorry. No. [00:50:53] **Alvaro Retana**: Still have three more. Okay. [00:50:57] **Sue Hares**: My question follows along Jay's actual question. He's talking about preprocessing my question based on the other statements that have been in this discussion today about AI. Have you done correlation statistics or regression statistics to see that your, that this information, which provides median and and various, averaging statistics, is actually correlated to slowdowns in particular RFC production? [00:51:39] **Jaime Jimenez**: Or No. Nothing. I haven't measured production of RFCC in any ways. But [00:51:43] **Sue Hares**: No. You're you're I think my term was [00:51:46] **Mark Nottingham**: Sorry. [00:51:46] **Sue Hares**: Imprecise. What you want out of the end of the process is good text in the appropriate time. [00:51:56] **Jaime Jimenez**: Right. [00:51:57] **Sue Hares**: In my statement there, I was implying the qual you know, the qualitative evaluation that says this set of inputs equals this set of outputs, both quantitatively and No. [00:52:11] **Jaime Jimenez**: I haven't I haven't done that. I haven't done that. [00:52:13] **Sue Hares**: Yes. I just wanted to be clear. It was a question of clarification. [00:52:18] **Jaime Jimenez**: Got it. Thank you. [00:52:20] **Colin Perkins**: Chris. [00:52:20] **Chris Box**: Hi. Chris Box. I'd looked at your website, and, yeah, it's great. You're offering the tools so people can just upload, identify the draft, and and it will run through. My question is, are you paying for that? [00:52:35] **Jaime Jimenez**: Yeah. [00:52:36] **Chris Box**: Okay. Is that sustainable? Because Yeah. It sounds like it's it sounds like it's valuable for the ITF, and we should [00:52:42] **Alvaro Retana**: Yeah. [00:52:42] **Jaime Jimenez**: Yeah. I'm more than happy to I mean, it's prototype. It's not I'm not selling anything to anybody. It's just I don't work in this. I I work in some other thing. Yeah. Yeah. I can yeah. You think I could make money on this? We can talk. [00:52:59] **Colin Perkins**: Oh, [00:53:05] **Jaime Jimenez**: I'm I'm cheaper than Mark. So but yeah. Yeah. The it's it's not expensive. I think there there was one slide on the cost. It's not We'll be [00:53:12] **Ignacio Castro**: passing the hat at the end of the session. Yeah. Priscilla. Hi. [00:53:16] **Priscilla**: Priscilla here. Just a quick comment. I have been using some of these models to check, the use of AI in scientific papers. And, usually, they are pretty bad when you try to check, like, a paper, like, from many years ago before AI. So if I know correctly, you didn't test with old RFCs. Right? [00:53:39] **Jaime Jimenez**: No. With not systematically. Yeah. But I did pass sometimes RFCs, and they rate lower. [00:53:44] **Priscilla**: Which was the results? [00:53:45] **Jaime Jimenez**: They rate better. [00:53:46] **Priscilla**: Oh, really? [00:53:47] **Jaime Jimenez**: Yep. I don't know why. I cannot I mean, I just know that the result was, like, in the below 6% and below than than the so the lower, the better. But yeah. I I I don't know. This in my opinion, the only thing that seems to be grieving giving ground truth is really the stylistic aspect, so improving that makes sense. [00:54:08] **Marcelo Santos**: Okay. [00:54:11] **Jaime Jimenez**: Yeah. All the boilerplate. All the it's just the text. But it could be that the training data I mean, I don't know. [00:54:21] **Ignacio Castro**: Thank you very much. So [00:54:24] **Alvaro Retana**: we're gonna go now into a non AI. Oh, a notetaker left. Yes. [00:54:33] **Ignacio Castro**: AI notetaker. [00:54:34] **Alvaro Retana**: So we're gonna need a backup notetaker. I know you raised your hand before. [00:54:39] **Jay Daley**: Somebody else could [00:54:40] **Ilke Ilhan**: help as well. [00:54:41] **Alvaro Retana**: Okay. There is an AI notetaker, but, you know, we want the human verification, all that stuff. As I said, we're gonna go into a non AI please stay. This has been, you know, a a very interesting topic, and I haven't heard Rasperjee mentioned so many times during the week as as it has been this week. And I know some of you came because it was mentioned somewhere else, and this is great. We are obviously a research group. We are not in the position to define policy for the ATF or anything like that. But, you know, we can be a place where we can talk about some of these things. There have been a lot of interesting questions, a lot of interesting ideas on what things could be done, what things could be experimented with, and we would like to see if there's interest in the room on maybe holding an interim towards, I don't know, September sometime maybe to continue the AI related discussion. To do that, of course, we need interest from you guys that you're attending, but also interest of people to talk about issues. Maybe questions that you can come up with in the list or somewhere so that we can seed a discussion. So I guess we should do a show of hands on the thing here. So the question is very simple. Would you be interested in a follow-up interim to discuss the AI impact on the ATF. Interim. Interim. If I can and it just says interested in an AI interim, so just please say yes or no. And we'll work with you guys on the list, assuming yes, to, you know, build an agenda and, you know, have people come and talk about this. Yes. [00:56:43] **Karuna**: Is there a reason to do something in September, like urgency as opposed to wait for the next idea in person? I'm just curious. [00:56:52] **Alvaro Retana**: No. There's no urgency. I've just I I I no. I think I just don't want to, you know, let the topic sort of lie down and say, let's wait till San Francisco. No. There's no urgency. And it doesn't have to be September. It can be, you know, whenever we figure it out. Okay. Thank you. We got some answers there. Thanks. [00:57:17] **Marcello Santos**: Okay. So I won't talk about AI. Please don't leave the room. Give me some support. So it's a initial step. Okay? I have more question than answer in this presentation. So we start to one simple question. Why draft is die or not? So like I said, we have more question than answers. So we collected around 35,000 of drafts from from the data tracker. And, unfortunately, we can use only 31 because we we don't have data in all of the drafts, so we can't track everything. And we have around 5,000 of RFCs [00:58:04] **Brian Trammell**: that we [00:58:05] **Marcelo Santos**: can correlate with this draft. So, maybe we can think about the conversion rate. Like, okay. We have 5,000 of RFCs and tier one, draft tier 1,000 of drafts. But when we get this draft is from the data track, they give the old draft is updated like version zero, one, two, three, four. So we need to we need to create a flow to understand exactly how this draft became RFC. So it's not the correct answer for this question. So quickly, the methodology, we collect the data from the data track, RFC editor, and we create a dataset. And like I said before, we have a little problem because we have an individual draft and many versions are replaced when they are adopted by our research group or our work group. And after that, became a RFC. But beyond that, we have some problem because sometimes two drafts became a RFC. So it's difficult to mapping the correct flow from a draft to RFCs. It's not so easy. So it we have here two example. Probably everybody know who's in this room. Internet draft can be adopted by a working group or research group and became a RFC. Or in some cases, individual draft can became a RFC without a working group or research group. So we have some stamps in the draft collected, expired, replaced, and published, active, and the dead. Some small portion are dead before the expiration, but only 200. And what we did was transform these drafts in a flow or lineage. So now we have a flow and we can detect if this draft and all the sessions that we count one time, once became RFC or not and some characteristics, some some features that, we can infer why this draft became RFC or not. So the first another question, how often draft became RFC? So now correct answer in the median, 21%. So but we need to look better about this result. It's not a direct correlation. We have data only from two and five in the data track. It's before, we don't have any, historically data about the drafts that we can track this correlation. So here we can see a peak in the beginning because they put a lot of drafts in the beginning of the year there. And it's not a direct correlation, but the number of drafts and RFCs during per year, each year. So, here, it's an interesting thing. If we take a look only about drafts adopted by a research group or a work group and the conversion in our RFC, we can see ETF has a higher conversion rate. And from 2000, 2018, we have this decrease. Why? I don't know. It's the data. We need needed to check. We understand better what happened here because we have a lower conversion from the draft to RFC. Of course, here, we have a better explanation because they didn't have time enough to become RFC because it's close to to, and it's very interesting. Like, I said before, we have around 35,000 of flows of drafts, unique drafts in a in a life cycle. 20,000 start, as an individual draft. 20,000. And what is interesting here? Around almost 3,000 are adopted by a working group and research group. And when it happens, 7% became IRFC. And and less than 1% became a RFC. So it's a strong signal, of course, for surprise of no one. If your draft is adopted by a research group or work group, they have more chance to became a RFC. But sometimes it happens that the draft born first in a work group or research group without individual draft. It's the graph in the middle. So when it happens, the draft born in a research group or work group, they have almost 71% of chance to became a RFC. But for my surprise at least, we have, almost 2,000 RFCs without any work group or research group adoption. And another question, if you check the IETF areas, what's the rate rate of conversion? So, of course, here, we have taken considerate take into consideration only the drafts adopted by a working group and the research by a working group here. So application real time have the higher conversion rate. General, the lower conversion rate. Why? I don't know if if the data. And in the amount of draft routing, of course, have the higher name, higher numbers of RFC. And in which moment in which month, we have the draft submitted. So we have this chart here. So, of course, close to the meetings. But here, it's October. It's not November. And I I started to begin. Why? Because November, it's the beginning, the first week. So October, we sent everything that we have to send. So, March, July, and October. What's interesting? Because we say many times, okay. We work online all the time, but, potentially, the data show that we we make a effort. And how long does it take for Internet to draft to become RFC? So in the median, two years and a half. But, we have a strange behavior after pandemia. It's almost one year more. Like, if you took the last four, five years, we have almost one year more to the draft became RFC. Why? I don't know. But it's interesting. This behavior may be with I don't know. We needed to to think about it. Here, another question. A draft became RFC faster in IETF or ER ERTF or IETF or IETF. It's another stream that we can send a draft. So ERTF take a little bit more longer. And which region has the highest number of Internet to drive the simulations? Take a guess. North America. Yes. North America. And the simulation and conversion, North American, if you took the whole data, has a higher number. And but it it's interesting. If we took the proper proportion. So we can see Latin American second place. Of course, we sent less drafts, but we convert almost 30%. So and it's also interesting. After the draft, is adopted by a work group or research group, it's almost the same rate of conversion independent of the region. But before Asia has the lowest, rate, 10%. And it's very interesting also. If you took the number of drafts and, submitted per year in the beginning of the data, North America, of course, it's in the top. But now Asia and Europe sent more drafts than than North America. Interesting. Here, it's a short picture about Latin America. We have Argentina and Brazil in the top. Let's see. Oh, I need to work. Here, it's also interesting. If we have a individual individual draft and we present in a presidential or remote meeting, we have 25% of the chance of the draft is adopted by the work group. We are a social group. So if you you if we defend our idea and present remotely or presentially, we have higher chance to to to to be adopted by the work group and here to become a a RFC. Another interesting thing, the number of the number of revisions of the draft. Of course, if you have more revisions, probably the community is engaged with this document and the higher probability to become a RFC also. If you if we have as a author, a director or a chair, of course, more experience, more chance to become a an FCO. So And if this draft go to the last call to become a FC, we have 93% of chance to become RFC, but we have 77% did not. So, my final presentation here, it's the end of the presentation. Imagine not all the people, but if you have a tool that we can check online quickly information about drafts and the groups, RFCs in a real time, like here, for instance, any NMRG with the numbers of the average time to become a FC, active drafts, regions, and so on. If you have a map, world map, where we we can click and see the drafts of these regions, and we have a range of the time, like, we if you can choose the year, the beginning, and the end. And see, I like my small BI to create some graphs and save. So we did that. We can access here. It's drafted to rfc.com.br. It's an initial version of this tool, of course, so we can prove that. So maybe we can discuss how to define what's a life cycle draft better and the conversion rate, how we consider that to to be comparable in this methodology. And, of course, every RFC has a history, every inefficient proposal that may hold a lesson. And that's it. Thank you. [01:11:38] **Ignacio Castro**: Thank you very much. Sue? [01:11:49] **Sue Hares**: First of all, thank you for your presentation and your work. This is important to see how we do it. This is a comment directed to methodology and not to the quality of your effort, but the actual methodology underneath it, which if I understood the last slide, were looking for comments and questions. Did I understand your last slide correctly? [01:12:15] **Marcelo Santos**: Yes, yes, yes. [01:12:16] **Sue Hares**: Here's the problem with your methodology. Based on a 10% detailed analysis I did of of drafts, meaning I actually did a hand analysis as well as a statistical analysis. The one thing that is caused that causes oftentimes changes in RFC production is after a working group does working group last call and sends it to the ISG. The ISG process has many loops which you are not accounting for in your methodology. You are taking an assumption that has been characteristic of many of the papers on ITF, presentation going back to 2000. I recommend this is a failed methodology because that process actually engages several pieces of the float. I and as you noted, the information is in the data tracker Yes. And you can pull it out. I recommend you look at that methodology, and I offer my services if you want to discuss the substates. I think once you do that, if you go back to your float and challenge, you will find why those years in RFC production caused some problems. The second part of your methodology, which I believe the recent, changes to the RFC production will give you additional piece is there are sub states in the RFC production which are not indicated in your draft. Your assumption is here's the draft, here's the working group, it gets adopted, it gets working group last called, goes to the ISG, and the rest of the process is a constant variable. I would again state that's a problematic assumption. Third point on your methodology. [01:14:15] **Ignacio Castro**: If briefly, please, because there are number of questions. [01:14:19] **Sue Hares**: Pardon? [01:14:19] **Ignacio Castro**: If briefly, please, because there are more there are more questions to have enough time. If you can make that third point briefly, [01:14:25] **Sue Hares**: set set list, then I will be done. But Okay. I am respond to working group last calls have problematic sub cycles. Just looking at all of these will improve your model. Again, kudos for working on this. [01:14:41] **Ilke Ilhan**: Thank you. [01:14:42] **Marcelo Santos**: Thank you. [01:14:44] **Colin Perkins**: Dirk? [01:14:45] **Dirk Kutscher**: Hi, Dirk. I wanna push back a little bit on, you know, treating ITF and IRTF in the same way. So in the IRTF, I don't think we should be so ROC fixated because there are many really useful contributions in the drafts that do not even want to be become an an ROC. And so I'm just a bit I mean, just a bit concerned if he sent this message that, you know, every everybody strives to become an ROC. That's not the case for for the IRTF, so let's just not equate this too much. [01:15:21] **Marcelo Santos**: No. Yeah. They agree. It's only a initial Sure. Sure. Stood, and we can debate and improve that. Yep. [01:15:29] **Jaime Jimenez**: Colleen. [01:15:29] **Colin Perkins**: Hi. Colin Perkins. Agree. This is a really interesting analysis, and I'm really glad you're doing this work. I would be interested to compare notes on how you find the draft history for an RFC in the data tracker. I I have done some similar analysis myself. My experience is that there are about four four or five different ways in which that history is recorded in the data tracker, all of which are inconsistent with each other. And resolving this is a bit of a challenge. And I don't think it's gonna significantly affect your results, but it would be interesting to see if the way I did this is the same as the way you did it. And if not, what we can learn from each other. [01:16:17] **Marcelo Santos**: Of course. I appreciate that. [01:16:20] **Jaime Jimenez**: Jaime. Hi. Thank you for the presentation. It's very interesting. On the draft submission deadline, like, there was one like, a graph with the bubbles and so on. I mean, I can explain that the work does happen regularly, but once you have the deadline of submission, you you tend to submit right before the deadline so you can present something. It's not that it's not work is not happening. But and, actually, I was wondering how can you see that. No. It was before. It doesn't matter. It was one of the slides. It's okay. How can you see that? And and maybe another data point for you would be the GitHub issues and the commits. Because now a lot of the work happens on GitHub, you can easily track it or the mailing list as well. Like, you can see that there is stuff going on as well. But, yeah, for sure, before submission, everybody is is sending. [01:17:04] **Alvaro Retana**: Okay. [01:17:04] **Jaime Jimenez**: And then you'd be interesting also to see rather than the country, which I assume is by you know, you can get these metrics, like, where the the origin of the person is or on so on. You realize the affiliation of the person and the headquarters of the company. That will be more representative as to the so you would have individual influence and then company influence and whether it has changed. That's pretty interesting stuff. [01:17:27] **Marcelo Santos**: Yes. We have a problem with the geography because sometimes we have this information Yep. From the auto. Sometimes we don't have. So what we did, it's not the best thing to do, but we took the email based on the email dot b r. It's from Brazil. But [01:17:45] **Jaime Jimenez**: Right. There's a lot of experts. I I I I I will be Finland, and I'm in Spanish. [01:17:50] **Marcelo Santos**: Yeah. That's right. Exactly. So it's approximated. [01:17:52] **Jaime Jimenez**: But it But very cool name is Nick. [01:17:59] **John Levine**: John? Hi. John. Again, allow me allow me to echo the comments. This is this is really interesting work. And I don't know how to qual qualify this, but somebody was saying, well, you know, drafts in the IRTF that didn't necessarily fail when they didn't join the DRR of seats. But draft drafts in the IETF, frequently the same thing. I've written a bunch of drafts where people sort of discussing something, and I say, I think this is a stupid idea. Yes. But I'm gonna write it up as a draft just so so so we have it written down. And then we talk about it for a minute, everybody says, yep. It's stupid. You know? And then it and then it goes away, which was was not a waste of time. We we needed to discuss it. You know? And I wish I could tell you some way to sort of identify drafts that were, you know, aimed for oblivion so you don't, you know, they don't count they don't count as failures. But at the moment, don't have anything brilliant to suggest. [01:18:45] **Alvaro Retana**: Mhmm. [01:18:48] **Ignacio Castro**: Thank you very much. And, from, draft mortality to community vitality, pun subtly intended. And we have, Ilke from, Ripe who has been doing a similar type of analysis, but, looking at the Ripe community and its mailing list. So over to you, Ilke. [01:19:12] **Ilke Ilhan**: Thank you. Hello, everyone. I'm Uke. I'm a business data analyst at Ripe NCC. And, yeah, I'm here to talk about the Ripe mailing list. This is a topic that is especially interesting lately because after decades, many people in the community are wondering whether communication might be mailing list might be getting irrelevant. Of course, big part of this is due to many other communications channels arising, but there is a bigger worry around the future of the community itself as many of the very active people in the community are getting older. So, young lists are the main communications channel for the ripe community and have been seen as a reflection of the community's vitality. So it's worth exploring these questions in the mailing list data, which is what I did, and I'm here to share my analysis with the IETF. So, this is an analysis on the right mailing list, but, of course, it's relevant for the IETF since these are two similar communities that's in a way grew alongside each other. They've been going on for around forty years. They've had, many, many mailing lists, a lot of them on, similar topics for similar purposes. But the similarity that I find more interesting is the traffic on their mailing list over time, which is what these two graphs are showing us right now. The graph on the right is the graph I took from paper on the ITF mailing list, which Ignacio was the coauthor of. It is the yellow line that I specifically wanna show you because that's the one showing the mailing list traffic over over time on the ITF mailing list, of course. And the the shape of that graph, looks similar to the orange line we have on the left for the right mailing list traffic over time. So both of these communities on their mailing list, the traffic, increased for quite some time, peaked around 2,000 tens, and then had a decline. I should say that I don't know much about the IETF, but I hope that as I share my findings on the right mailing list, it will make you think whether any of it applies to the IETF. And before moving further, I'd like to share my scope. So these are the right mailing lists that, I have in scope for this analysis. They're the most popular right mailing list that made it into the twenty twenties. Most of them are working groups and community mailing list with the exception of members discussed, which is exclusive to the membership, and right path list, which is dedicated to the Internet measurement tool. So it is this scope that adds up to form this graph that we just saw briefly. We said this shape was like there was a long increase, a peak, and then the decline. But to me, this kinda looks like the shape of an existential crisis. The decline at the end after the peak in February make some people in the community wonder whether communication might be dying off. But as I'm touching on such bleak themes here, I'd like to remind everyone that this isn't over. Graphs tend to reveal their true shape in retrospect. So this could be dying off, but this could just be a setback before surging up even higher. Or much less dramatically, it could just be stabilizing around this level. We can't know for sure, but we can't what what we can do is to look at the state of it now. So let's focus on the twenty twenties and look at the current situation of mailing list participation. And I think this is a really interesting question to start off with because it this is a question that gets asked a lot a lot in the community, is engagements driven by just a handful of people. And this is a concentration graph that responds to that question. So we have all the accounts that sent an email to one of these mailing lists in the 20 on the x axis, sorted from the most active to the least active, and the y axis counts the running some of the emails sent, which, of course, reaches a 100% in the end. Let's make this a bit easier to read, and we get a very clear answer that 19% of the contributors sent 80% of the emails in the twenty twenties. So this is a high con concentration, but it's also common enough to have its own name, the Pareto principle or the eighty twenty rule. But I don't wanna stop here. Let's divide these accounts into five groups that each sent 20% of the emails, which would look something like this. A lot of thin bars on the left, so sorry. Let's focus on them. So at the left most side of this graph, the top top contributors, we have eight accounts that sent 20% of the emails in the 20 They are followed by groups that sent the second 20%, the third, the fourth 20%. And overall, there are around 1,600 accounts that make up this graph. And this is a high concentration even though a common one. So does this suggest narrow participation or that there's a dedicated core? And since we are contemplating on the longevity of these mailing lists, are these top contributors old? But I should say that we don't have the data on anyone's real age, but what I can do is put these accounts into generation bins based on their first email to one of these lists, which is how we're going to proceed in the following graphs. So this graph has the same 1,600 people divided into the same five bands that each sends 20% of the emails. This time they're, colored by the generation they're in based on the first email they sent. So for example, we have the people from the nineteen nineties all the way in the bottom, shown in yellow. And on top of them, we have people from each of the following five year periods stacked on top of each other with the people from the twenty twenties all the way at the top. The people from the twenty twenties, in fact, make a big part of this graph, which is already a great sign. That means that we have a lot of newcomers in the twenty twenties. But it might not be fair to have them in this comparison since many of them likely joined very well into the twenty twenties, which is our scope. So I'd like to exclude them and compare the other generations. So on the left side, when we look at the high contributors, it seems that, high contributor groups are more likely to, be from the older generations. But not all high contributor groups have the same profile, so I'd like to focus on them separately. On the left side, we have the two thin bars that has, that have the people who sent 40% of the emails in the twenty twenties. And almost half of them have been contributing since the nineteen nineties or February. So if you're on one of these mailing lists and you're familiar with some of the people in the community, a lot of the emails that you receive come from old timers. But these are outliers, and there's a whole other picture, when we look at the other side of this graph. Now I'm looking at the groups that sent the third and fourth 20% of the emails, and they do seem more likely to be from the older generations, but less so. It looks a bit more balanced. So beyond the few very high contributor groups that we have, there is a larger, moderately high contributor group that is not at all dominated by the older generations. In fact, this whole graph shows us that there is quite a lot of younger or newer contributors in the twenty twenties. But let's zoom out of our twenty twenties focus and see how it's been historically and whether the community has been renewing itself. So this time, we have the periods on the columns, and we still have the generation split through the colors. So the people from the nineteen nineties, the emails in the nineteen nineties, they were all sent by people from the nineteen nineties. If you look at the graph on the left, which counts the the accounts that sent an email in each period, and the right one counts the emails that were sent. So people from the nineteen nineties in yellow, few of them survived into the next period, the early two thousands, and much fewer survived all the way into the twenty twenties. The same can be said about the following generations as well. Not many of each generation stayed active for long enough to be counted in the next five year period. But their the contribution of those who survived has been significant if you look at the graph on the right side. It seems we're getting the most out of those who stuck around in the long term. The other side of the story is, of course, the newcomers, and it seems from the graph on the left that the majority of contributors in each of these periods periods are newcomers, and newcomers sends more emails than any other generation before them. But it's still definitely not an equal contribution. When we look at both graphs, we see that the surviving oldies are sending more more emails than the newcomers that we get, which is, in my opinion, natural that the newcomers might be feeling things out of it before starting contributing heavily. But still, I'm curious at this point to see the long term contributors in each of these generations. So let's highlight them. Now we are focusing on the accounts that stayed active for at least ten years. There are not that many of them if you look at the graph on the left, but they have sent many of the emails as we can see from the graph on the right. But the more interesting thing is that they didn't necessarily start fast. They didn't immediately start sending so many email. And in fact, their contribution somewhat correlates with the overall traffic we see here. So it seems, they sent, they talked more when there was more to talk about, which brings me to my next point. We wonder who comes to the community, who says who contributes, but why people would come, stay, and contribute is very much about what is discussed, what is important, and what is going on, which can be observed through the different mailing lists we have. So let's go back to scope to remember what we're dealing with. We already know that these are the most popular right mailing list lists, but let's see how they compare to one another. So these graphs show us the number of emails each of these mailing lists have received. The first one shows us all the emails all time, and we have a very clear winner. Address policy has received the most number of emails all time by far. But the picture is quite different when you look at the graph on the right, which only counts the emails that were sent in the in the twenties. Now we have a new lead, members discussed, followed by Ripe Atlas. This is such a big shift, and the ranking of mailing list shifted in every period as different topics attracted people to join, which is what this colorful mix of visualization shows us. We can observe the journey of each of these mailing lists as they through their changing relative popularity. Among all, I'd like to focus on at risk policy and members discuss. At risk policy, it seems, immediately got the top spot after it was launched in 2003 and stayed there until the twenty twenties when members discussed crept up to take its spot at the top. It makes perfect sense because when we look at when address policy was this busy, it was the time of IP for run out and the challenges that came came with that, which the working group had to respond to. And after the run out with, not much more resources to allocate and the shrinking membership, of course, the focus shifted to the charging scheme, which keeps members discussed aflame. So it shows us that traffic is very much topical and also dependent on external factors. But how much traffic each mailing list has is one thing. It's another to see which ones bring more newcomers since we're talking about the longevity of mailing list. So this graph has shows us the number of first time emailers each mailing list has brought. And very clearly, members discussed at the top brought the most fresh blood followed by Ripatlus and addressed policy. But it's one thing to bring the fresh blood, another to see whether that fresh blood stayed in circulation. So now we see that those accounts divided by whether they, sent another an email to another mailing list later on, which is shown by the dark blue parts of the same bar. So if we look at the dark blue, it seems address policy takes the lead again. It has the most number of newcomers that later contributed to another mailing list. And when we look again to members discussed, right, Atlas, it seems they don't have as much of that dark blue. Most people, newcomers that first came to discuss membership or write Atlas, they don't wanna jump to other topics to other mailing lists. This makes some sense because they're kind of different as we highlighted in the beginning. But also considering that these two are the most popular two most popular mailing lists of the twenty twenties, it does make one ask whether the community is in fact going silent as the members discuss and the Atlas community talks. So it's not that straightforward, but I'd like to go back to this total that we started with. But this time, let's see how much each of these mailing lists contributed to the overall shape, and let's start with address policy again. The shape of traffic on address policy seems somewhat similar to the overall graph we have, and the peak is, mostly driven by address policy. And we know that address policy had that peak during the time of IPv4 runouts and all the challenges that the community had to respond to. So, it makes us see this overall peak under different lights. Perhaps what it shows us is not a time of peak community engagement, but a time of peak challenges for the community, which the community responded to. After that period, we had members discussed as the most popular mailing list, but its traffic is not as consistent, which makes sense. We know that members discussed tends to explode on occasion, mostly as a reaction to discussions around the charging scheme or the NCC's budget. So understandable. But Atlas, on the other hand, is the opposite. It has consistent high traffic, which, again, makes sense because it is much less dependent on external factors. So yeah. So these three are very different mailing lists. All lists are different. Different working groups have their own dynamics. So I would say simply comparing their numbers doesn't make much sense without the added context. It is the insights we get by analyzing the meaningless data beyond just the total number of emails sent over time that can help us interpret the shape that we started off with. So, yes, we have the decline after the peak in, February. But throughout this presentation, we've seen that most of the contributors in each of these periods are newcomers. And even though they're not making much of the noise just yet, those who stick around are likely to talk more in the future. Today's long term contributors, we saw that they didn't start fast. And in fact, their, contribution increased when there was more to talk about because people come for topics. People don't engage with the community because they're a part of the community. They follow discussions they find important, which makes them a community. And mailing lists provide a platform for these different topics to be discussed on. They all have their own dynamics, so their traffic is not really comparable. And since their traffic also depends a lot on external factors, even the total number they adapt to doesn't translate into the community's vitality, which is how I interpret this shape of an existential crisis the mailing list data has plotted so far. I hope it made you think about the IETF. If you have a different interpretation or questions, comments [01:37:02] **Ignacio Castro**: yeah. Thank you very much. Peter? [01:37:08] **Peter Koch**: Yeah. Thank you. Peter Koch, Dienyk, I'm on almost all of the mailing list that you mentioned, actually. And in your accounts, I'm a person from the nineties, and I I wish that was true or not. I think the remark that you made towards the end that the numbers might spur interest but don't make sense without the without looking into the quality or the content is very important. And that's probably also where the outliers, may or may not have anything comparable in the ITF environment. If you look at the Atlas list, for example, the usual traffic is a researcher, usually a young researcher, having a project asking for Atlas credit points and then thereby generating 10 responses where people will shortly report, here's my credit, and thank you, and goodbye. And the respondents will usually be old timers or older timers and reoccur whereas that initial message from the reporter will not researcher will not be followed up. And that's, consistent as in it's happening over and over again, and that's most of the traffic there. Same for other working groups. If you look at the content, it's like the ITF. Right? Doesn't vote, but everybody likes to count the number of interventions. And, so you see a flood of plus one emails that is also very bursty traffic. And, again, that is something that may or may not occur in the IDF or be comparable. But I think, again, that remark is important, and here's some anecdotal evidence to what what's happening there. I don't think it supports the existential crisis, but it absolutely deserves further insight. So please continue going into that into these measurements and also looking into the content of the messages. Thank you. [01:39:07] **Ilke Ilhan**: Thank you. Hi. [01:39:12] **John Levine**: I'm John Levine. I have a it's it's I mean, those were some very interesting charts. I have a a mechanical question, and then I have an anecdote. And mechanical question is, did you have did have you been able to track people when their email address has changed? Because it appears to me there are not a lot of people yeah. I mean, I'm wondering how many of those newcomers are just different addresses. In the ITF, we have a so so way to tell that multiple basically, if you've ever if you've ever submitted a draft, you're in a database and and we attempt to combine all your email addresses. You know? But depending on how much extra work you wanna do, you might wanna go back and look at the the, like, the the name on the front line and see whether, like, if if if an address disappears and another one reappears with the same name on the front line. Maybe maybe it's not a newcomer. Maybe it's just an old old old timer who got a new job. Okay. The anecdote is I happened to have lunch today with a bunch of grad students at the local university, and they were talking about the work they were doing and all this other stuff. And one of them said, yeah. I once tried to join a mailing list, but it didn't work. You know? And it was, you know, it was like it was supposed to be the the help mailing list at his university, but it turned out it had been set up ten years ago when it was sort of on autopilot, nobody was reading it and writing it. So there's definitely you know? And I and I said, no. Really, there are mailing lists that work. You know? And and and just because one is dead doesn't mean they're all dead. But, I mean, I but I think we definitely have a generational issue that, like, know, people who look like me think that mailing lists are normal, and people who look more like you or some of the other younger people here think that, well, you know, it's like it's kind of this quaint old thing that that that we do to do to work here. But, I mean, there's there's so many people that say, well, we do you know, we use Discord and we use GitHub and, like, I suppose, a mailing list. So I think the existential issue is real, but I don't think it's insoluble. [01:40:56] **Ilke Ilhan**: Thank you. [01:41:01] **Colin Perkins**: Hi. Colin Perkins. Really interesting analysis. I think what would be a really interesting thing to do would be to run a sentiment analysis tool on some of these messages and see if those bursty lists are all just people yelling or if it's other types of discussions as well. [01:41:23] **Ilke Ilhan**: Good input. Thank you. [01:41:27] **Colin Perkins**: I think we have a bunch of ATF lists where it's all just people yelling. So I hope I hope you're different. [01:41:37] **Ignacio Castro**: Nick? [01:41:39] **Speaker 2**: Yeah. I think thanks for this work. I think we have seen some some similar attempts of this analysis analysis at ATF. It it already got mentioned some, but I am curious what you think about, like, displacement to other communications for the this is always the discussion that comes it up with ITF or w three c or others is well, maybe the mailing list traffic has just changed because GitHub has taken off or people are using IRC or Slack or whatever. And, you know, like, that that's not gonna be, some answerable conclusive question. But but I am just already curious, like, how do we track other communications for, may maybe we could even see it in the mailing list data or something. But, like, if if if we are trying to measure activity or or vitality or or tenure of participation or things like that, how should we be looking at multiple communications media? [01:42:35] **Ilke Ilhan**: Yeah. I mean, that's a big question. Mailing lists are certainly one aspect of participation or vitality. So, yeah, of course, there is more. Yeah. That's something to think about. [01:42:50] **Ignacio Castro**: Thank you very much, Elke. [01:42:51] **Ilke Ilhan**: Thank you. [01:42:55] **Ignacio Castro**: And for the last presentation, we have colleague Perkins with a draft and license Internet standards. And just for clarity, I'm also a call for here, so I could refrain to my co chair. And let me pass it control. Use this. Yes. In a sec. [01:43:25] **Alvaro Retana**: It should work now. [01:43:26] **Andrew Campling**: Yep. Okay. [01:43:29] **Colin Perkins**: So hi, everybody. My name is Colin Perkins from the University of Glasgow. This is joint work with Ignacio, with Rio, who is sitting over there, and with Steven McQuiston from the University of Saint Andrews. We have spent a reasonable amount of the last few years working on doing this sort of analysis of the ITF data. Looking at the the the the mailing lists, looking at the history of the documents, looking at the history of and and changes in the participation. And we submitted this document to try and help structure our thinking about how to do that systematically And to try and help provide inputs to the research group, try and take some of the lessons we've learned from learning figuring out how to do this. However, well, we'll have a badly even we're doing this. And try and pass that experience on and help other people learn from those lessons and and structure their process for for doing this analysis. We spent you know, we we submitted the draft. It looks at the standards development process in general as a sociotechnical system for for taking input producing documents. It looks about some of the issues of doing this analysis in terms of [01:44:51] **Jay Daley**: the [01:44:51] **Colin Perkins**: ITF, how we might apply that more broadly to other standards development organizations, some of the challenges in data processing, some of the issues with ethics, data protection, and so on, and then tries to make some recommendations. And it's very much a work in progress. We're looking for input to see whether people agree that this makes sense and whether it is useful. So we start by considering the standards development process, as I say, as a sociotechnical system. Status development is a process of taking a bunch of inputs. You know, there's some technical artifacts, Internet drafts, for example, mailing list messages, software. This various input from people, from organizations, you're reflecting their organizational interests and priorities, and from the the governance process of the standards development organization, which sets the rules by which everything works. That runs into the standards process, which produces some some you know, eventually produces some standards, eventually produces some implementations of those standards. And there's various feedback loops. The process of developing the standards iterates a bunch of times, and you get feedback from the standards and from their implementations back into the whole thing. We try to characterize the type of data we're working with. Right? We've clearly got participants. We've clearly got organ organizations and and, you know, the, you know, the the organizations those participants work for, but also other organizations in the community which might be affecting the way things are developed. Your governments, for example, civil society organizations. We've got various technical groups, working groups, research groups, areas, directorates, and so on in the IETF. We've got the artifacts, you know, the drafts, the the the presentations, the the mailing list messages, and so on. We've got the actual collaboration infrastructure. Is it a bunch of mailing lists? Is it meet echo? Is it the Java room and so on? The various different ways in which people communicate. We've got the governance structures and processes, you know, the the ITF rules, the the way in which the ITF operates, the other SDS operates. Then we've got the standards and implementations that result from all of that. There's a lot of things you can automatically extract from this. There's a lot of things you cannot automatically extract about the process. One of the things that is very clear from doing this analysis is the type is that the type of data we can extract automatically by looking at the data tracker, by looking at the mailing list, by looking at the the conversations and so on, is obviously useful evidence, but only captures a very small part of the standards process. And a whole bunch of really critical aspects of the the the developments are really hard to observe direct directly and in in some cases impossible to observe directly. There's obviously a culture to the standards development body. There's influence that peep particular individuals, particular organizations have in that body. Does the agenda setting process that the influence and who sets the agenda, both formally and informally, and who's who's having influence in the organization. And anyone who's worked in the ITF for a while or in any other standards body will have a feel for that. They'll know who is influential. They'll know whose opinions matter and who is, you know, a newcomer who's just turned up and and no one knows whether they're making any sense or not. And that changes over time as people gain experience, as they submit doc documents, as they show that they understand what's going on. But this is not visible externally. And while you can capture some of that by looking at measures of communication, for example, you you can capture some measures of influence. They're clearly only very approximate. Similarly, informal coordination negotiation, who exercising power, who's exercising authority, and how is very hard to see from the data. So we can only capture a part of the process. And the data we can capture, the different metrics, different artifacts, and so on is very different in terms of accuracy, relevance, degree to which it is representative. If we try to turn this into looking at the ITF, there's various data sources we can we can get access to. The obvious one is the data tracker, which has the metadata about the process going back to the early two thousands with varying degrees of accuracy. We've got the mail archive. We've got the RFC editor of websites with the RFCs, the errata, the RFC index, and so on. And increasingly, we've got GitHub, and we've got things like the the video recordings on YouTube and and the chat logs and so on. The ICF makes a lot of data available. If you just pull the data tracker, there's about five and a half million records in there. If you pull the mailing lists, there's 3,000,000 emails, give or take. That's about 40 gigabytes of data. That's before we pull any of the RFCs or the Internet drafts or the presentation slides or the video recordings or anything from GitHub. So there's an awful lot of data to work with. Considering the other STOs, what they make a very what they make available and how they operate varies significantly. I think the the ITF is probably one of the most, if not the most open in terms of data. But certainly, the the you know, all these other organizations make some data available. And if you're a member, there's a lot more data which is available quite commonly. So there's a lot of data you can collect for the other organizations. Their goals, their moods of operation, the way they work, the way the the way the data should be interpreted is very different across the different organizations. And integrating the data and and interpreting it it across the different organizations is a significant challenge. Things to worry about when looking at this data. Some of the challenges that that we ran into. Entity resolution is an enormous challenge, both tracking people and the organizations and tracking changes in affiliation. But certainly in the ITF and I think in a lot of other standards to own organizations, there are people who participate for a very long time and significantly you have a significant influence and change their names and affiliations all the time. I think I appear in the ITF data tracker under about four different variants of my name. There are other people who are their names are expressed in many different ways and possibly even who have changed their names. And tracking that, especially if it's not well recorded in the data tracker is a significant challenge. Similarly, organizational names. The previous version of the slides I had, I think it was what, 280 something different spellings of Huawei that are in the data tracker. [01:52:24] **Alvaro Retana**: That's And [01:52:26] **Colin Perkins**: that's that's easy to track to that's easy because they all have Huawei in them. But, you know, tracking organizational names and how and the the different ways in which people express their organizational name. Tracking mergers and acquisitions of companies is, you know, a a significant challenge if you wanna track who's working for whom and tracking organizational influence. This is also really significantly complicated by consulting relationships, which tend to be very, very opaque and by various opaque funding sources. And as an academic who gets research grants from all sorts of people, including industry, tracking who's funding me to be here at any particular time is perhaps a challenge. So it's a put putting all this together, just figuring out who's talking, who they're representing is is very, very difficult. Susan mentioned this earlier, but the document life cycles are difficult to track. In theory, this is in the data tracker. For the ITF, the data is wildly inconsistent. And for any particular rule, if if you if you if you read RFC twenty twenty six and all the ITF process documents, for any rule you can find in those documents, there are some documents in the in the history which violate it wildly. There is not a straightforward process of individual draft to working group draft to send it to the ISG, and then it gets published as an RFC. Things cycle at all all the different points in this this process in in very interesting ways and bounce between IETF and IRTF in different groups and and who knows what. Leadership role history is a challenge to to to reconstruct mostly just because the data tracker is awful. It just really does not track who to who is in which role in terms of working group chairs, area directors, and so on very, very accurately reconstructing this is hard. If you try to analyze the mail archive and you think this is easy, has a bunch of codes to manage email messages, we can just throw the IMAP code through the Python IMAP library at the ITF mail archive. It does not work. There are so many malformed messages which break the Python library, especially as you go back in time. And you have to do a significant amount of cleanup of the data to get it working in common email processing tools. The number of emails have two two addresses in the front line, for example, which just break all sorts of things. There's thousands of them. And the the historic data is just wildly incomplete. ITF has good records for about the last twenty years, but earlier than that, it's very, very inconsistent. Ethics, privacy, and so on is is a big challenge. Obviously, the ITF has a privacy statement and a bunch of policies for access to data processing this. This is all personal data, which falls under the constraints of things like the GDPR in in different parts of the world. Your research ethics committee, your institutional review board, whoever it is, are going to have opinions on this and probably should have opinions on this as is your data protection officer for for whatever organization you work for. So some interesting challenges there. I would also say be careful what you release. Right? This data is all public, but the implications of it are not necessarily obvious or well known. We did some analysis of RFC errata a few years back, for example. And we can we can tell you who writes the RFCs, which have the most errata against them. That may be sensitive. We can tell you who is the most successful at writing Internet drafts or who is the least successful at writing Internet drafts if you analyze this data. We're not publishing that, but this sort of data is available if you hunt. So you have to be careful what you release. It's it's potentially sensitive. It potentially affects people's careers. And this, I guess, comes to the issue with AI. The effective function functioning of the standards process affects critical infrastructure. We don't wanna break the organization by releasing a bunch of information that will cause a meltdown and, you know, everyone to start fighting. You know? So what are our recommendations? I don't have time to go through all these in detail. I don't think I have space to go through all of these in detail. But the draft makes a number of recommendations both to the ITF and to researchers. To the ITF, it talks about the need to preserve the data. You know, preserve stable access, data quality, take care when backfilling historical data to make sure it's done correctly, take care about the provenance of the data. So we know what's original and what's derived, and if it's derived, what it's derived from and how. Managing the process changes. The the what's in the data tracker varies significantly as the data tracker has been improved over the years, but also the process has changed over time. And it's really hard to to know exactly which sets of rules the organization was working under at any particular time. And that can sometimes affect the results. We make a number of recommendations for the researchers. Many of them are fairly straightforward data science recommendations. Take care with the data tracker. It's very, very hard to do this right. And there's zero dot zero useful documentation. The those those ways which look obvious of extracting certain types of data, which are just wildly wrong and miss important things. And there's no way of there's no public documentation of how to extract particular types of information, which involve reconstruction. Identity affiliation data, we don't have good ways of doing this. I don't think we have an agreed community consensus way of reconstructing who's whom and what who do they work for. And that that means results are very difficult to compare across papers because everyone's doing the reconstruction in a different way. Be careful interpreting the metrics, engage with the community, try and try and release code, release approaches, and maybe RESTBudgy can help here. There's it it's hard to get this right, so let's make sure make sure we we we share experiences. So that's broadly it. There's obviously lots of different ways of studying standards process of studying internet governance. The approach we take is clearly only one of them. And we've seen talks I mean, Karen Kaf has given a bunch of talks in the IRTF about an ethnographic approach to studying this, which is a completely different way of studying the organization. You I'm not saying this way is is better or worse. This that you you learn different things from different approaches. If people ask are following this approach, do people find this type of documents, this type of guidance useful? If so, is it something Raspbatch should be considering for publication? If not, well, hopefully hopefully you find at least this talk interesting. [02:00:07] **Alvaro Retana**: Thank you, Colin. Brian, so we're at time, but we can, of course, stay here longer. So but keep your comments. [02:00:15] **Brian Trammell**: First question, yes. Second question, yes. One of the things you've identified with this draft is that the data tracker is basically meant as a system for running the operations of the IETF, but it is also de facto. It's archived. Yes. That might be a question to raise with, I guess, the tools team via the IASG. I would recommend that you do that because that is actually important to understand that as a requirement. Yep. Thank you. [02:00:49] **Alvaro Retana**: Priscilla? [02:00:52] **Priscilla**: Hi, Colleen. Hi. Following Ryan, yes and yes for the two questions. I have some comments. Please. Strike for me. It's it's it's you are right. It's really wrap it's really hard take all this data from Datatech. I'm doing this for the last two years, and I I feel your pain. Yeah. It's really hard. Especially, after saw all the presentations today, I started to think that something that I'm doing maybe is helpful. I don't know if it's a good thing to add in your your draft. I read the draft last week. It's pretty good start. But one one thing that I'm doing is, I'm taking the data and doing analysis. And then sometimes, okay. I only have data from this year. And sometimes, oh, I don't have data enough. This week, I just figured out that some people are have more than one country in the data track. So sometimes, like, yeah. Almost 4,000 people change the countries during the meetings. Yep. So it's it's how how to how to deal with these situations, you know, not just these, countries, but also you say they change their names and how to deal with. So I start to write like a guide for myself and from people that work with me. So I'm put things like, okay. This data, I only have from this year. Oh, this data has this problem and creating, like, putting suggestions about how to solve these problems and how to deal with them. Like, kind of guide that they can use working my projects. And I also doing your presentation, I was thinking about there are some AI conference or data science conference or venues that they have, like, a checklist that you can check if you do the preprocessing all the things correctly. And maybe this is useful for people that are working with ITF data. Yep. Because many of the meetings of this research group, there is someone that oh, this is not the right way to do the preprocessing of this data. This is not the right way to clean this specific data. And this happened because there are people here that are working with this for long years or here, like, for years and know why this data is missing or whatever. Yep. Yep. So maybe it's good for us discuss this too, like, create, like, a checklist of things that you need to think, and how to deal with the data. Not exactly this this, but suggestions about how to read and put, like, the most common misleadings, for example. Yep. Maybe this is can be really useful for the group. [02:04:01] **Colin Perkins**: Yep. Yep. Yeah. I mean, I I think this sort of whether it's code or just checklist, this sort of guidance for things to think about when doing this analysis would be really useful. Yes. [02:04:11] **Priscilla**: Yeah. Yeah. Exactly. And especially when people come here to present the analysis, you'll be useful is the last thing. Maybe use it for people also say, okay. This was the way that I did the preprocessing. This is the limitation that I found. So that's it. [02:04:28] **Colin Perkins**: Absolutely. Yeah. I don't know what the best way of capturing this is, but I think it'd be really useful to capture this. [02:04:32] **Alvaro Retana**: Five [02:04:33] **Jaime Jimenez**: seconds. Five seconds. A lot of the people gathering stuff and data use the API calls, the the data tracker. I find they are sync way better, and I my feedback would be, like, nobody's handling that. Just populate it with more data, like, and so on. The are sync is the right tool in my opinion. [02:04:49] **Colin Perkins**: Yeah. I mean, some of the data is only available in the data tracker, but extract extracting it in an incremental way is extremely hard. So you have to end up pulling several copies, and they they the tools team just get annoyed with you. Yeah. [02:05:03] **Alvaro Retana**: Before everyone runs away thank you. We don't get too many requests for adoption or publication or anything else. There seems to be a lot of interest. We're take this, obviously, to the list. The question in the in in the slides and everything was about considering publication. We'll, of course, consider that later on. Of course, we'll first talk about adoption. Thank you. [02:05:25] **Colin Perkins**: Thank you. Thank you, everybody.