Markdown Version

Session Date/Time: 22 Jul 2026 09:30

[00:00:40] Marcus Ihlar: Alright. It is 11:30 on a Wednesday. We are in Vienna and it means that we we're starting the first of two IPPM and BMW joint sessions this week. Welcome, everybody. Yeah. This is an IETF meeting so please note well, we're in the middle of the week so you probably have seen this one a number of times. But please familiarize yourself with the processes for contributing to IETF meetings regarding IPRs, behaviour, etcetera, etcetera. The note well is here. If you're not familiar, please take a look at it. Note really well, this is a professional meeting. We are here to have professional and polite discussions. It's important to behave well and treat each other with respect. There are guidelines around codes of conduct and anti harassment procedures. If you feel that you are not being treated well, there are several ways to handle that. So yeah, we have an onboots team and so on that you can contact. So please note this really well as well. Right. Some meeting tips. Please make sure if you're in the room, sure that you have signed into the session. This is very important for planning for future sessions around slices of rooms and so on. You can use the QR codes, which are found by the microphones and usually outside the room as well. Or you can join from your laptop using the on-site tool. If you're joining remote, please always keep your audio and video off unless you are actually speaking. It's strongly recommended to use a headset for audio quality purposes. This is a joint meeting between IPPM and BMWG. We're still two working groups, so we still have two independent charters. As usual, we just quickly show the two charters here. Here's a snapshot of the BMWG charter. The important aspects of this is that it's around key performance characteristics and benchmarks of devices and systems, and typically not operational networks. IPPM,

[00:03:41] Jared Mauch: on

[00:03:41] Marcus Ihlar: the other hand, is focusing on operational networks, so it develops and maintains standard metrics that can be applied to the quality, performance and reliability of Internet data, etcetera, etcetera. You can read more if you're interested. We have a note taker so we can skip this. Thank you very much, Ike. Today's agenda, this is the first of two meetings. So we will start off with updates on working group documents. We have a short update on the Connectivity Monitoring metric, which is a document that has been parked for some time. We have update from Ike on quality of experience. This is a document that has already been progressed out of the working group but we're getting an update on the current status. And then Stuart will be showing some hackathon reports and other things around IPPM responsiveness. There will be a few presentations on proposed work as well. We will have three presentations on IPPM and BMWG topics, pretty much. Once we have done that, we will have a joint charter discussion. There has been work going on between the previous meeting and now on a joint charter. There has been some good discussions and Ching will be presenting the outcome of these discussions and we will have some questions to ask around this as well. Does anybody want to bash the agenda?

[00:05:33] Ching Wu: To anyone who enters the room, it would be

[00:05:35] Marcus Ihlar: great if you close the door because they don't seem to close themselves, unfortunately.

[00:05:41] Ching Wu: You.

[00:05:44] Marcus Ihlar: Alright. A little bit about document updates. Do you want to talk about the BMWG updates?

[00:05:49] Giuseppe Fioccola: Yeah. For BMWG document updates, we basically submit the segment routing benchmarking methodology to ICG for publication. We received the the review from Matt, thanks to Matt, and also awaiting reviews from RTGWG and performance monitoring there. In addition, other two documents are in RFC editor queue, so the young model and the MLR search.

[00:06:22] Marcus Ihlar: Alright. In IPPM, we have quite some updates. We have this connectivity monitoring, which is a parked document that has been parked for a while but is on the agenda today. Quality of outcome is now in the RFC editor's queue so that is a great milestone already and it's on the agenda today as well. We have hybrid two step that was progressed some time ago and it has received AD review and the authors are working on addressing that. We have stamp extension header where we had a working group last call recently which has now been concluded and we're waiting for a Shepard write up. Waiting for me actually. But yeah, hopefully that will be concluded relatively soon. And we have two documents that are still in working group last call. It's been for a while, it's Altmark-deployment and Altmark-yang. These have received really good and extensive feedback and Giuseppe is working on that. We hope to conclude this last call not too far future either. Asymmetrical packets is in RFC editor's queue as well, so that's very nice to see. And then we have IOAM data integrity under ISG evaluation and we see very good progress there as well. The integrity yang has also been submitted to ISG for publication. So we have been draining our queue of documents pretty well and we hope to keep that up. So thanks to everybody who contributes to this. Right. Giuseppe, do you want to talk a little bit about 99 bis?

[00:08:00] Giuseppe Fioccola: Yeah. This topic was already discussed on the main list. We got the support from the working group, you may recall, to move the RFC seventy seven ninety nine as a well established document for IPPM as a BCP, best common practice. But at the last call, the proposal received some mixed feedback, especially to suggest to refresh the document. And for this reason, I published the RFC seven seven nine nine. Please, we already have some discussion on The Middle East. We want to move this document forward quickly. So review and comment, we also have the GitHub page for this document. So yeah.

[00:08:56] Marcus Ihlar: All right. And that brings us to the end of our chair slide and we'll move on to the presentations. So we will start with connectivity monitoring.

[00:09:18] Rüdiger Geib: Oh, thank you. Okay. Assume it's just one slide. Thanks very much. I'm with Deutsche Telekom. Yeah. Experimental connectivity monitoring metric has been parked for a while. It's segment routed related, and, yeah, I want to finish this work before I finish my professional career. And made the trusted way forward. It's to add a definition of success criteria for this experiment, and then this document could proceed. And that is what I did. It's the first draft now of this condition. First thing to do is, of course, to set up a lab environment or also do that in a live network following the conditions given in the document. The idea there is that by segment routing, you can specify the path of measurement. And if you create a specific overlay, you can evaluate the results by some tomography methods, and from that, take conclusions on the status of the network. Where where is the connectivity lost? Is congestion appearing? Well, no, not just that it is happening, but also where it is happening. That means there's certain simplicity in all that. It shouldn't be extended over the globe. It's rather in local networks, but it can be done reliably. And the document is not about routing. I have been asked that several times. It is about a bit longer terms, which are on the base of some several seconds and up to minutes, but it is faster than the usual ten to fifteen minute counters which you use in operation and data to figure out what the network does. Right. The three criteria are mainly as whether the round trip delay can be determined for each monitoring link as defined by the document. There's a definition given, how you can evaluate the measurements specified there to determine the round trip time of each link which is monitored. And then the next one is whether the loss of a monitored, a single loss that's very important, the single loss of a monitored link can be detected as defined by the document and whether the congestion of a single monitor node interface can be detected as defined by the document. Now, these are the most important metrics and specifications in there. And, of course, it would be great if these results are published. Well, yes. And if it works as expected, then I think the experiment would be concluded successfully. Otherwise, we can drop it. Feedback is welcome, of course, to prepare for our last call. As an experimental RFC, it's quite important to state it's experimental because there is no implementation yet. Unfortunately, it has been planned, but it's not there due to events which I have no have no influence upon. Are there any questions? If not, thanks.

[00:12:54] Marcus Ihlar: Tim is in the queue.

[00:13:03] Tim Chown: Tim. No problem with the document at all. Just something is a knit for the chairs. If you look in the data tracker, it's still listed as intended proposed standard rather than intended experimental. In the document, it says experimental. But just to keep the ADs happy, probably change that.

[00:13:20] Marcus Ihlar: Thank you for that. We'll make sure to update it. Yeah. Thank you so much for picking this up again. It's highly appreciated. And I mean, this document has been parked for a while so I would, before going to a working group last call, I would urge the working group to actually take a look at the document first and provide reviews. And, yeah,

[00:13:40] Participant: we'll take it Yeah. From

[00:13:59] Ike Kunze: Okay. Hi, everyone. So I am Ike. I, end of last year, joined the QoO draft as an editor and would like to present a short update today. So as Marks has already said, the draft is now in the RSC editor queue, so that is very nice to see. And we also recently have been seeing growing and growing interest across a lot of different dimensions. So we now have two to three implementations of quality of outcome. We are seeing large scale rollouts and field trials with Tier one network operators. And we're also seeing a lot of research and development projects with academia, with content providers, operators, and with Windows. So there's really a lot of interest that is currently growing. And the main reason why I asked for the slot here today is to make sure that we actually keep quality of outcome interoperable. And why is this actually a challenge, or why am I bringing this up? Because the quality of outcome framework itself is intentionally defined to be very flexible so that we can deploy it in a lot of different scenarios, which also means that if a lot of operators or vendors make their own choices, it can become quite the implementations could diverge. When I talk about deployment choices, what specifically am I talking about? So just a brief summary or a reminder of what quality of outcome has as its main inputs. On the one hand side, we have network measurements. There is a lot of different ways you can measure the underlying raw quality of service metrics using active measurements, passive measurements and all that in a lot of different ways. On the other hand, we have per application requirements that we need for quality of outcome, where we have latency percentiles, where we could use different percentiles. We have loss rates that we also need to measure and throughput. And for each of those, we have an unacceptable and an optimal threshold. So also a lot of different ways how we can configure those. Which means, ideally, we would like to have convergence on those different configuration aspects so that when network operator one says we see a very high quality of outcome score for all of our customers, and operator B has just slightly more ambitious thresholds, which is why they are seeing slightly worse quality of outcome scores, then ideally we would like to prevent that. And so what I would like to start as a discussion, not necessarily only today, but essentially in the upcoming months, is about sharing deployment experience among the group here and interested parties to compare and maybe also think about how we can actually come up with different application profiles and what kind of thresholds we choose, what kind of latency percentiles we choose. Then based on that shared experience, also start thinking about documenting common best practices, and then also identify which of the areas you might not have really captured in the original quality of outcome framework, so that we can then also augment that afterwards. What I started during the hackathon already is to develop a small demo where we can experiment with different application profiles, different kinds of measurements in a very simplified docker based setup where you can essentially emulate different network configurations and then just try out what different application profiles will have an impact on the scoring and also what different measurement approaches will tell you about the current network state. So that is really just the start to get a discussion going. And what I would like to do next, as I already said, is to see if we can or would like to also, at some point, bring this back here into the working group in the form of documents that give these best practices or give these recommendations and maybe also standardise the way that or standardise specific application profiles. So there's a lot of different topics that could be of interest for here. And with this talk, I essentially only wanted to raise awareness for that and will also post to the mailing list after the session or later this week so that we can then get started there. And what I would like to be or what I'm specifically interested in is to also get more and more application developers into the picture to also come up with sensible application requirements. There are a couple of definitions out there already, but they are, I think, mostly coming from operators' views. And ideally, if we really want quality of outcome to work well, the application requirements should also come from the applications themselves. And that is all. Thanks.

[00:19:02] Marcus Ihlar: Thank you. Do we have any questions? Brief ones?

[00:19:12] Participant: Thank you. Very supportive of this work. I think I agree that it's very important to have industry involvement And also, some granularity is needed such as for gaming, we're doing also the pro gamer community because they have this very strict requirement. So we need to develop also very precise profiles. You.

[00:19:41] Jason Livinggood: Hi. Jason Livinggood from Comcast. I really like the dashboard that you showed earlier and and all the work that's happening. And we are planning to do some field testing of this stuff soon, so get some operational experience. And as usual for all these things, bring some of that feedback here. Thanks.

[00:19:59] Marcus Ihlar: It would be very interesting to hear about it in this group. Thank you so much. Thanks, Ike. And Stuart is up next.

[00:20:18] Stuart Cheshire: I'm Stuart Cheshire from Apple. We have a full agenda today, so I'll make this quick. But this is a working group document, so I owe the working group a status update on what is happening. I refreshed the draft with a few minor changes before the cutoff deadline, so draft zero nine is available. And I spent this weekend at the hackathon working with Greg White and others. There was a student joined us, Hamid Hassani. I wanna thank him for his help. And we were looking at the problems that reported. Spent the first day on Saturday getting very frustrated. You can see this pretty scary packet trace here. The the white packet line is pulling away from the act line. The queue is getting deeper. The round trip delay is going up. There's all the purplish sack blocks, the registry transmissions. I was pulling my hair out. And then we found out that there was a kernel version mismatch with the TC tool on the bottleneck that Greg was using. So, thankfully, we had not broken TCP Prague in OS. That would have been very scary. So we kinda wasted a day struggle with struggling with that, but at least it wasn't work worse. Once we got it sorted out on Sunday, just got some beautiful traces of L4S and Prague doing exactly what it's supposed to. If I zoom in a little bit on this, you can see there's a steady stream of packets. There isn't a single loss or retransmission in the whole flow. The round trip delay is kept bounded. The simulation we had was a 20 megabit bottleneck with ten milliseconds of delay in each direction. And here, the round trip delay is the base twenty milliseconds plus one or two. So it's pretty consistently twenty one, twenty two milliseconds. Anytime the queue started to creep up, we'd get the CE congestion marks. The sender would make a small measured adjustment to its rate, get it back under control. The queue would shrink, the CE marks would stop. Just a joy to see it working correctly. However, the responsiveness RPM tool was reporting two hundred and sixty milliseconds of round trip time. When clear from the packet trace, it was 2122. I spent some time looking at the code. Unfortunately, I could not figure out what was causing it, and I don't wanna speculate right now because I literally don't know. So we clearly have some more work to do. I know we wanna wrap up this draft, but we don't wanna publish it when it's producing results that aren't useful. So I was in Slack communication with my colleagues at Apple. We're gonna get some of the same hardware that Greg White was using, and we will continue to get to the bottom of this and figure out something is wrong in the test methodology, and I don't know what it is, but we will find it and we'll fix it.

[00:23:29] Marcus Ihlar: So a quick question without my chair hat on. Do you believe that this is a specification issue or an implementation issue?

[00:23:41] Stuart Cheshire: I think it's a bit of both. The Apple implementation of the tool was written by Christophe Pausch before he left Apple, and he's also a co author on the draft. So the implementation and the draft are pretty well in alignment. Right.

[00:23:58] Marcus Ihlar: Then do you have any other implementations to test with?

[00:24:02] Stuart Cheshire: I did not test at the hackathon, but yes. Is Hawkins, the professor at the University of Cincinnati, has an implementation in Go, and there are a couple of others that have been AI generated from just giving them the draft and saying implement this. So, yeah, we have a plan of the work we need to do. I my I said I didn't want to speculate, but my current guess is we we really wanted a test that was very quick that you could run it, and in ten seconds, it would give you a score for your network. And with the time it takes congestion control algorithms to converge to a steady state, maybe we have to relax that goal a little bit, that it it is gonna take more than ten seconds to get a repeatable result. Because the critical thing is you need to be able to run this test and get the same answer every time. If it's like spinning a roulette wheel and it's kind of random what answer you get, then it's not a useful test. So Christophe Pash really wanted it to be a quick test, and and he was influential in the draft. And I think maybe we're gonna have to walk that back a little bit and be pragmatic.

[00:25:27] Marcus Ihlar: Thank you. Stuart, you have a short announcement to make on We thought we'll optimize here and let you do that. Well, you're still up. Sorry.

[00:25:44] Stuart Cheshire: So another quick one. I'm presenting this on behalf of Bob McMahon. You may recognize that name. Bob is the maintainer of the iperf-two performance tool. You've probably also heard of iperf-three, do not be led astray. Iperf-three is not better than iperf-two. It's a pity that they didn't just pick a different name for their tool. It's so everybody naturally assumes that iPerf-three is the replacement for iPerf-two, and it's not. It's a different tool with a different focus and different strengths. That's kind of unfortunate confusion. IPerf-two is alive and well, and Bob, with the help of some AI tools, has taken his core protocol analysis engine and wrapped it with an Android GUI app, which gives you wonderful real time displays. My favorite one, because I care about latency, not just throughput, is it used some Linux kernel calls to show you the state of the send buffer. So it'll show you the currency win value, the bytes in flight. It'll show you how much of the buffer is data that has not even been sent yet. It's waiting for the congestion control algorithm to give it permission to send that data out. Of course, all that adds up to delay from the application layer perspective between calling send and the data actually leaving the machine and traveling through the network and arriving at the other end. And seeing that update live as you're running the test is great. There are some screenshots here. I'm not gonna dwell on these slides. You all know how to go to the agenda and download the PDF. So if you're interested, grab it yourself. It's only for Android right now. There is not an iOS version. If you go to sourceforge.net, there are a bunch of precompiled APK files. Just pick the right one, you can install it on your Android phone and send feedback to Bob or on the mailing list. I think it's a wonderful tool. I need to get an Android phone myself so I can run it now.

[00:27:52] Marcus Ihlar: Ask your employer. Alright. Thank you.

[00:28:10] Alexander Clemm: Hey. Hello, everybody? Can you hear me?

[00:28:13] Marcus Ihlar: Yes. Hear you fine.

[00:28:15] Alexander Clemm: Okay. Excellent. Alright. Okay. Yeah. So I'm going to report on in the in-network telemetry aggregation using two IOAM options, which was also part of a hackathon report. So this is kinda like a combined presentation on what happened at the hackathon as well as on those two drafts. Accordingly, there's a joint work with a lot of people. You see their names listed here. Next slide, please.

[00:28:41] Ike Kunze: You you have the the the controls, Alex.

[00:28:45] Alexander Clemm: Oh, how how do I do that?

[00:28:48] Ike Kunze: You should see at the bottom of

[00:28:50] Giuseppe Fioccola: the slide.

[00:28:50] Marcus Ihlar: Yeah. At the bottom of the slide,

[00:28:51] Alexander Clemm: should see this optional arrow. Okay. I think I got it. Thank you. Alright. So the background well, basically okay. I don't think I need to explain IOAM to this audience. So, clearly, basically, this concern is collecting telemetry data from nodes across the network path. The thing or the aspect that sets this work apart is that we are looking specifically at aggregation traces. So rather than having well, recording and and collecting the individual records from each node, we would preprocess it during packet reversal using some aggregation function. For example, minimum, maximum, sum, and so forth. There are a couple of use cases for this. Actually, our first use case was was also with network sustainability where we wanted to assess carbon metrics across across the path, but there are other use cases as well when you plan to find bottlenecks or outliers on the path and and and and what have you. So the basic idea is to pick it down there. You have basically the nodes. So you have the package with the IOAM data traversing the network, and, basically, every node updates the aggregate as opposed to adding its own records. So alright. Yeah. Let's go to the next slide. So the starting point was this aggregation trace option alternative, and there's a there has been a draft for this for a while. Actually, since then, basically, there's also now a second draft on a template. It turns out that you can use also templates to to express through the same thing with an aggregation trace template, and the fields actually align, yeah, very well. In the interest of time, let's go forward. So which brings me to the hackathon. So the the goals of the hackathon was really basically two folds. One thing, basically, we want to actually, yeah, really compare these options in detail and extend the existing proof of concepts to to to make sure that all of these are covered. There are two tools of concept here. One was basically surrounding this network sustainability case. It's basically maintained by, yeah, by the OST University in Switzerland. And there's a second one which basically extends the Linux kernel implementation with support for these options, which is tempered by the University of Rome. And we wanted to also well, the question came up about the use cases and and whether it would make sense to also not just report the aggregate, but also record the nodes that were actually traversed. And so that was basically kind of, like, a goal that we pursued at the hackathon. Can we combine these can we combine those two? Obviously, they could be handled using separate IOAM options, but there are some issues with that. So instead, we basically collapse them into another combined IOAM aggregation option which we call the trace plus template option. K. Let me move on. So you see, basically, the new option defined here. It say, basically, okay. I won't bother you with the individual fields, but, basically, the the upshot is that having this option saves eight of overhead per packet versus having the separate options, basically one for parity, which is significant since we are talking about attaching this to production traffic. Likewise, basically, in the course of this, we updated the template and the template option. And because it turns out there are a couple of bits that we can actually save that aligns things better and makes things a little bit cleaner. And this is documented now in the template option, the new revision zero three, which we just posted yesterday, which reflects, basically, also the outcomes from the hackathon. So this one here is just a very brief just to give you an impression of one of the two POCs. This is basically the one with use case with this network well, with the sustainability of of network paths. The idea is that you have a certain efficiency indicators, which are, yeah, which basically give you information concerning the carbon intensity at any node at any one particular moment in time, and you basically add them up. And the idea is that, basically, by looking by by collecting, aggregating this data, you can distinguish between paths which are cleaner and which are not so clean. So they basically this is done on p four running on VMP two and, basically, on the on on on on the Ubuntu Linux machine. There's a monitoring stack that's that's also attached to this where we collect the exported data and ultimately, basically, just display it. And you see, basically, here on the top, a, yeah, basically, a heat map or other things that support it can be generated from that data as the packets get reversed. Obviously, you cannot see this, but basically, on the on the left hand side, obviously, you see the Wireshark output here, so you can actually see this is actually IOAM data aggregation data that is used to produce these results. And it's like there are, in fact, actually two implementations, but this is the screenshots of of one of them. Alright. So in terms of other updates, actually, we also updated the other draft, the aggregation trace option, again, reflecting some slight optimizations to the fields that we made as the outcome of the hackathon. And, yeah, essentially, well, what we would like to ask from the working group is, first of all, actually, for the IPPM, IOAM aggregation option as well as the template option have been around for a while. And I think we've reached the the point where it is where we should ask the working group whether it wants to go forward with this. Basically, if we can go to towards adoption in the foreseeable future, that is basically, yeah, the the first set of ask. And the second ask is basically the question if there's interest in the trace of template in this combined option, which covers the use case where you want to have an aggregate as well as basically the, yeah, recording information about the the notes on the path. This pretty much concludes what I have. In the backside, you see basically the the references to the draft, actually, as well as the links to the POCs so that you can play with it yourself if you have. Okay. Thanks. Question?

[00:35:40] Giuseppe Fioccola: Yeah. Alex, I have a a general question. So the work, in my opinion, is interesting as a contributor, but as, let's say, chair of BMW, I'm interesting whether since there are several IOAM extensions document, I'm wondering whether you checked with the the authors of these extensions if they would like to be aligned with this template in the future if adopted? So something like that. So did did you check did you talk with the authors of the other IOAM extensions that are currently proposed in IPPM?

[00:36:19] Alexander Clemm: I'm not sure with which one you you're referencing right now specifically. The as far as the aggregation is concerned, we are aware basically just of the IOAM aggregation draft as well as in the template, of course. And we are collaborating. And, actually, now all the authors are collaborating together to see if we can align them. Actually, I think the the the consensus is that should the template option go forward, the template basically allows you to define the structure in the template that you don't have to repeat. Yeah.

[00:36:50] Giuseppe Fioccola: That's why I'm wondering whether the the the other people which are working or other IOAM extension can start to be aligned to this IOAM template. So this is the the

[00:37:04] Alexander Clemm: So so of well, the general answer is of of course, yes. But I'm I must say, actually, I'm I'm I'm personally not aware of other work that would be overlapping with this. If you look at the

[00:37:14] Giuseppe Fioccola: IPPM documents, there are some I'm extension just to check and to start to maybe it's a good time to start to contact the authors of these I'm extensions and and check them to, you know, to to involve more people and get them aligned with this IOAM template. So

[00:37:34] Alexander Clemm: we'll yeah. We okay. We'll we'll do that. And if you are aware of somebody we should specifically talk to, also please let me know.

[00:37:42] Ching Wu: Okay.

[00:37:46] Giuseppe Fioccola: Maybe Frank?

[00:37:47] Frank Brockners: Yeah. This is Frank. So, yeah, first of all, thanks for doing all the work, and the hackathon results were very nice. So we've been doing that template option basically to do exactly what you said prior. Right? So we try to go and harmonize rather than coming up with new options all the time, like harmonize the format of how we would do that, which is why we created that template option and Telesat that, like, maybe it's the last real option that we do, and then we just fit that in and we do something very similar to what we've done with IPFIX. Right? Remember, we started off with specific fields and then we came up with templates so that the whole thing became way more streamlined. And we're trying to go and do the exact same thing with IOAM. So maybe it's about time to consider working group adoption for the template thing so that we don't really see that option sprawl moving forward and at least have a foundation that we can go and work on as a working group.

[00:38:50] Marcus Ihlar: That sounds good. We would like to just do a quick check to see how many people in the room have a and we should probably do it for the separate documents so we can start. Yeah. And here we're talking about the template option then. So please, any clarification question? Do we have someone in the queue? Otherwise, the queue is locked.

[00:39:19] Alexander Clemm: I just had one more response perhaps to what Frank was saying. We are yeah. So completely agree with the template offer. We don't want to avoid the sprawl. However, some of the templates, because they carry semantics, the fields, they they may still need to be subjected to to to their own draft as well. It's pretty, actually, the aggregation stuff, unless we want to collapse collapse those things.

[00:39:46] Marcus Ihlar: Okay. I think we're starting to see some numbers here. So we can record it. Okay. Yeah. There's still some more coming in here. Yep. It seems like around seven people have read it. It would be nice with more, but it's not nothing. So thanks for that. This is a useful signal and we'll take this into consideration. Yeah. Okay.

[00:40:20] Ching Wu: Good morning. And my name is Ching Wu, and I'll join with other calls are to present this AI fabric benchmark suit. So we have three draft terminology, training, benchmarking, inference benchmarking. The current version number is zero three. So a little bit of recap for the document status. Actually, the zero level version has been first presented in the central meeting, and we got a lot of broader review and discussion. And thanks, Carson and Matt and Colin and many other actually to provide good input and suggestion. And, also, we have a new and and. Yeah. Actually, he we based on his, you know, input, actually, we actually invite him to add a. So we created GitHub to track all the open issue for this story draft. And, actually, just before this meeting, we actually resolved a lot of the open issue and with PR. And there so what changes since previous ITM meeting? Actually, for terminology, we try to better, you know, align with training, benchmarking, inference, benchmarking, and for terminology. Wanna centralize the terminology and move, you know, more common terminology to the terminology draft and only leave some, you know, technology specific terminology in the training benchmarking and inference benchmarking draft. In addition, we add, you know, some missing terminology acronym in appendix a and b, and also we add a new term we call the packet radar improvement. This apply to the more like, you know, optical UET layer. Actually, we can use some, you know, compression mechanism to really, you know, reduce the packet overhead. So this is some new term we add. And, also, we make some, you know, correction for some definition in the terminology jobs. For example, bus bandwidth and roof line job completion time, actually, formula, which try to make it a precise reflect the way already, you know, implement. And and also for zero impact for failover, we also, you know, redefine and make it, you know, outcome based on not a limited specific mechanism. And for training benchmarking and also inference benchmarking, actually, we also make a lot of update. For example, we added additional load balance configure. We support UEC reliable on order delivery, support dynamic and load up flow net. And also, we one issue is about, you know, UEC port port port, actually, whether we need to assign a port, you know, in this draft. And also, we remove one reference, which is, say, twenty eight eight nine. And because not many term we, you know, use this, so we downgrade it as an informative reference. Inference, the benchmarking, we also, you know, make an update, make some of, like, error number more illustrative, not normative. And we break it down the good port into the inference good port and a good port. And so this relation with the existing benchmarking draft that we clarify. You know, we try to build on top of the, you know, benchmarking work already published it. And and, also, we, you know, clarify our relationship with the UEC and documents. So this because we want to tackle another just, you know, RoCEv2 RDMA, but also we want to deal with u UET transport. And so this, you know, next step, actually, we think, you know, we want to people can keep an eye on some KPI definition test procedure, and we will we think this is already a good starting point, and we make a lot of, you know, updated this. We would like we can consider adopt this kind of work. So I see.

[00:44:28] Marcus Ihlar: Please be brief because we are short on time. Yeah.

[00:44:31] Participant: Hey. Hi, I'm from University and have two of these questions. And the first is that where do you think that the benchmark should be operate on in a tester, a specific tester, or separately in the real servers? And the second question is that what do think is key differences between the testing and operating online? Yeah. Thank you.

[00:44:54] Ching Wu: Yeah. We mostly focus on, you know, library environment testing, but we do have planning to cover some, you know, more production like, you know, environment tester. So this should be considered. Yeah. Thanks for your comments. We can have more discussion offline about it.

[00:45:11] Participant: Okay. Thank you. Thank you.

[00:45:13] Ching Wu: Yeah. With that, I conclude. So have you considered my question? Yeah.

[00:45:21] Alexander Clemm: For sure. Yeah.

[00:45:22] Giuseppe Fioccola: I'm I'm one of the quarter, but anyway, we already discussed with the other BMW chair, Sarah, that this is one of the document candidate for the group of document candidate for the adoption. Yeah. Yeah. But we will further discuss with Markus and Thomas.

[00:45:40] Ching Wu: Yeah. Thank you.

[00:45:51] Marcus Ihlar: Right. And we're a little bit behind on the agenda. So if you can be a little bit brief, it's good. Thank you.

[00:45:56] CATS Benchmarking Presenter: I'll to save some time. Yeah. Okay. So this is a CATS benchmarking methodology. So CATS is a traffic engineer approach to for you to, I mean, steer the traffic to appropriate computing instance. So this draft is it definitely is about, like, defining some methodologies for CATS approach. So it has been the fourth revision. So we have received several good comments from the list and have incorporated them into the draft. So authors and always seek coordination between cast working group and BMW previously on technical discussion on the list because this document was proposed in BMW, but I think it's for all forced into the scope of the potential new merged working group. So quick summary about the revision here. So previously, before version four, the the overall document has approaching mature on including methodologies for, like, three CATS deployment modes, five core benchmarking testing to covering CATS metrics, collection, distribution, and some key performing indicator, like session continuity and end to end latency and overall load balancing in Verizon, and then reporting format accordingly. And in the latest revision in version zero four, we have added the test test setup for CATS hybrid mode and refinement on the mapping between test cases and their corresponding matrix with based on the recent updates on the test matrix framework draft. Improvement and also improvement on the expected results of each test case, and we have added some recommended metric aggregation functions and normalization functions in appendix. So we're thinking it is approaching mature. So next step so I have, you know, talked to, like, Joseph whether this can be adopted in the list. So I'm thinking it definitely falls into the scope of the potential new working group because it has some, you know, work on the measurement of the cast metrics and and some new test methodologies. So we respectively request that if the working group can consider the adoption of this document. Thank you.

[00:48:17] Giuseppe Fioccola: Yeah. As I also mentioned in our discussion, this is one of the document where we have also the expertise outside the BMWG working group.

[00:48:27] Participant: In this case, it's CAT's working group, but you are good to get

[00:48:30] Giuseppe Fioccola: a lot of feedback from CAT's experts. So, yeah, this is also another document that will be considered for the adoption soon.

[00:48:38] Participant: Thank you.

[00:48:40] Marcus Ihlar: Thank you very much. And let's change it up again. And now we will be discussing the charter of a possibly joint working group here.

[00:48:50] Ching Wu: Okay. This is the last topic. And so I want to give a quick update on IPPM and BMWG joint joint charter proposal since, you know, before this meeting, we have a lot of discussion and prepare the, you know, the charter update, actually. So we can merge it for the discussion. So happy to report the the status and the summary. So just a background recap, actually, this, you know, already discussed in the meeting. Actually, we conducted with many sending experts in this area and to summarize what is the key concern, challenge, benefit. And so based on discussion, the general agreement is take option two, like, you know, merging or target to one single charter. So in during the session meeting, actually, we also have follow-up, you know, join a charter discussion and open to everybody. Actually, we have Carson and also Luis and many other to join this and provide a good input. Thanks for that. And we also continue the discussion on joining charter proposal after Shenzhen meetings. So we try to make it a better, you know, policy to simplify the charter proposal. So you can see the charter proposal link here. And so for initial charter proposal, we, you know, kick off this discussion in April, and we try to summarize, you know, what are key elements we add. For example, we're not limited to the, you know, the protocol at a specific layer. Actually, we document a common security building block. And we also, you know, produce a terminology, recommendation. This, you know, more apply to benchmarking work. And, also, we cover some management aspect for them, a young data model and net effect net network can match it, such as API fix. And we also can cover extend to support a broader range of applications and encourage more collaboration within ITF with many working group and also liaison with many STO. So we receive a lot of input and suggestion, and we actually key change we add if we don't make distinction between the production and the library environment applicability. And, also, we, you know, reference some, you know, transport related based on the transport a a d coverage comments. We cover the scalability aspect, and and you you also can see, actually, we discuss the implementation policy and also the working group name. So I and I will you know? So one, you know, important fundamental change we had added here is the IPFIX for the management aspect. So and also system calibration for test and the metric makes it more clear and based on the comments we received. So thanks to many people to contribution. Actually, we finally to get, you know, more concrete of China proposal. And so two issue I wanna discuss here. One is about, you know, implementation policy. And this, actually, we find open issue on the data tracker on the GitHub. And the three implementation policy has been, you know, investigated. Now option one, we prioritize document with implementation. Option two, we require at least one invitation for standard track document. Option three, we actually raise the high bar, and we require at least two implementation policy for standard track document, very similar to what the IDR is doing. Based on the discussion, actually, we see the general agreement should take option two. So we, you know, to incorporate these changes, you know you know, with the merging PR. So you can see the linker is here. Second is about, you know, working group name. Actually, we also want to choose the right name for this joiner worker. Actually, several name proposal has been proposed. So here, I want to ask, you know, share maybe Paul for each name. We have five name choice. Which one will be best? You wanna do this? Yeah.

[00:52:38] Marcus Ihlar: So we're gonna run a little name poll here. So we actually have one know already. Is this person willing to come up and state why you're not happy with the name?

[00:53:27] Tim Chown: Hi there, It just sounds like something per minute. BPM, beats per minute, something per minute. It just we've got Stuart's RPM. Maybe there's a there's a thing. It just sounds yeah. Not like something I'd use. That's all. Sorry.

[00:53:45] Giuseppe Fioccola: It may also record the bits per minute of the network. Right? Something like that. Somehow. We are 70.

[00:55:56] Marcus Ihlar: Alright.

[00:56:06] CATS Benchmarking Presenter: That's the last choice.

[00:56:08] Giuseppe Fioccola: Yeah.

[00:56:09] Marcus Ihlar: Yeah. I think we're okay. Yep. So, yeah, it's settling in here. It seems like we at least have one that is the least disliked. Yeah. We'll take this and we'll probably take it to the list as well. And, yeah, Of course, if you have very, very strong opinions about this, it's good to speak up. But given that no one else speak up spoke up on on why they didn't like the other names, it seems like the opinions are not that strong. That's probably good. Yep. We have we have Jared in the queue.

[00:57:06] Jared Mauch: Yeah. Yeah. I just wondered, which of the two, a or b, would be for the BPM.

[00:57:14] Marcus Ihlar: Do you wanna have a new poll? Yeah. Maybe that's something we could take to the list as well. Okay.

[00:57:20] Jared Mauch: Let's, yeah, let's take that to the list. Yeah.

[00:57:25] Marcus Ihlar: Okay. But it looks pretty much like BPM or BIP are the, like, yeah, probably most popular here. But, yeah, let's confirm on the list, and thanks for the input.

[00:57:35] Ching Wu: Okay. Thanks, everybody. Very important. Yeah. We move forward. Actually, this next step, I think we ready actually to make it ready for the ISG review. And I think there's some housekeeping, you know, to have working group calls review. And and also after this, IAG review, maybe we also need to, you know, send a liaison to relevant and standard body. So any comments for this?

[00:58:02] Marcus Ihlar: Yeah. Now is a good time to voice any concerns?

[00:58:15] Tim Chown: Tim again. Yeah, I've already, as you know, given you some feedback on the charter. I think there's three or four things that could be improved. I think one is around the disruptiveness of the measurements. It says it will minimize the disruption, but I think we should say it may be that we define things that are temporarily disruptive. I think that's okay. The second one was around security building blocks. I think that's not clear what that is. Maybe just remove that bit because if you took out about 10 words there, it may may be more sense and be less confusing. Yeah. So I think there's the three or four little bits in the text that where it can be improved slightly. But overall, I think it's pretty much there.

[00:58:55] Marcus Ihlar: Thank you. And would you possibly mind writing up some GitHub issues or PRs on these topics? That would be very helpful.

[00:59:01] Tim Chown: Yeah. Sure.

[00:59:02] Marcus Ihlar: Yeah, that yeah. Thank you. That goes to everyone. Like, if you read the charter and you you see something that you're you know, want clarified or whatever, please, contribute with issues as well. It helps us drive this forward. Okay. Thank you very much. We're at the end. We'll see you again tomorrow for a longer session. So don't miss out. Thank you.

[00:59:31] Tim Chown: So we're have four chairs.

[00:59:36] Marcus Ihlar: Well? Yeah. Excited. Yeah. We well, actually, we have three chairs, and maybe we can play this game. You know?


Session Date/Time: 23 Jul 2026 07:00

[00:00:27] Thomas Graf: Alright.

[00:00:36] Marcus Ihlar: It's 09:00, so we'll slowly get started here. Welcome to the second session of this week for the joint session of BMWG and IPPM, possibly the last time we operate under these names. We'll see. We will confirm, but it's likely that a future name will be, BPM. But let's see on the list. Yeah. Do we have the click thing? Nope. Well, it's Thursday. Most of you have seen this a number of times, but, yeah, please be familiar around how we operate at the IETf, IPRs, etcetera, anti harassment policies and such. If this is new to you, please scan this QR code and read a bunch of documents. Note really, really well that, we conduct this in a professional manner. If you feel that you're not being fairly treated or feeling that you're harassed in any way, there are a number of procedures. There is an ombuds team that can help you in these situations. You have the resources here as well. Meeting tips. You should be aware of these. Please make sure that everyone in the room signs into the session. It's very important for us for planning. Join the mic queue before you go up to the mic. If we have show hands tools, you use the tools specifically for that. If you're not using the light client, please keep your audio and video off. If you're remote, also keep your audio and video off unless you're chairing or presenting. And as usual, it's a good idea to use a headset if you are, remote when talking. We do have a notetaker. Thank you very much, Stuart. We have a very packed agenda today. A lot of interesting topics, and we also urge everyone to, yeah, try to to to stay on time as good as possible. We will start with a suite of documents by Giuseppe, where a number of these have been are in working group last call, but there are, yeah, there are relations to to more documents. So Giuseppe will do an overview, and then Nalini will talk about an update to encrypted PDM. We will have a PowerBench presentation, SEB benchmarking, and we will have IOM for MPLS. So that is adjacent work to the IPPM, I would say. Then Greg will talk about class of service ECN for stamp. We also have a number of presentations on proposed work, as you can see here, ranging a large amount of topics. I'll not go through them all here now. If possible, if we make it to this slot, we have a number of lightning talks. The idea with these lightning talks, these are all if time permits. If we make it here, we would urge anyone who presents to do one slide, brief overview of what it is you're working on so that we can go through all of these. Thank you. Is there anybody who wants to bash this agenda? Take that as a no. Then we move forward with the first presentation from, Giuseppe.

[00:04:40] Giuseppe Fioccola: Yeah. Okay. Good morning. This is an update about this family with set set of documents about the alternate marking. So this is the first one about the deployment framework that has already concluded the working loop plus code recently, but I just we just want to provide an update to give the full picture and to see how the different documents are interconnected. So the first document about the deployment explain how to deploy the full solution and especially it clarify different aspects related to the applicability scenarios, like the controller domain, the rule of the nodes, the type of measurements. But in addition to that, the main point is that the whole framework also define the different options for the configuration, so the possibility to use YANG, PCEP, PGP, and for that export, so YANG push and IP fix. And the reference, the relevant documents that are also in other working groups. So, for example, the young model is in IPPM, but for IP fix is in OPSAWG. So and we are discussing with the OPSAWG chair where to do the to do a joint working group plus call for this document, for example. After the working group plus call, we received some feedback. The the new version address several comments from Benoit, Alex, team, and all the comments have been addressed. So we especially clarify in the introduction that this document is the base document. There are other complementary documents. We also explained that there are two options for the computation. The the first option that there is a collector. The other option at that considering the recent publication of the RFC ninety nine fifty one that each node can do directly the computation on the node assuming that the counters and timestamps are carried with the packet. You can see RFC ninety nine fifty one for reference. We also extended definition of controller domain to consider the case of trust relationship between different entities. Of course, we added the operational consideration section as well. For the young data model, we also update the young model and the the main change. Thanks again to Benoit and Alex for the reviews, we clarified that this young model is only for configuration. There is another young model that is for telemetry, and it collect telemetry for not only for alternate marking, but also for IOM and on path delay. We redefined the model to include the profiles under the interfaces because it's more coherent. And we also revised all the cross references between the corresponding parameters in YANG and IPFIX just to help the implementers to understand the correlation between the parameters across the different tools. For the push, similarly the same. We also, in this case, we addressed the comment from the ASI, Benoit, and Alex. And we redefined the model similarly to the to the configuration model, and, also, we add cross references and the operational consideration as well. We are still waiting for the performance metrics here review from Paul that will be also made jointly with the configuration model in order in order to check the consistency between the two. The last one is in WG is the IP fix extension. Just a recap. Basically, we use the same information elements that are already existing, but we need to add new information elements to accommodate alternate marking needs. So in particular, the flow mode ID, loss flag, delay flag, and the period number. Benoit also detailed review this document, and the the relevant changes are explained here. So we improved the description of your flow aggregation, correlation, and measurements. We changed the name to align with the naming convention of IP fix. Also, in this case, we clarified the two option for computations because after the publication on the RFC ninety nine fifty one and the ninety nine forty seven, there was another option for computation that is now enabled. And, of course, we checked the cross references. That's all. Next steps. So the deployment draft and the young model just concluded working group plus call. So maybe for the deployment, it can be moved forward. For the two young model, we still need the review to check, as I said, the consistency. The young will be updated soon in preparation of the working plus call. And, also, for the IP fix, there is a plan to do a joint working group plus call with Opswap WG. Of course, comments are work. Comments?

[00:10:31] Marcus Ihlar: Alright then. Thanks a lot, Giuseppe. Now p d m v two. Do we have nothing here? Yes.

[00:10:58] Nalini Elkins: Okay. So some of the comments that we had had on p d m v two had to do with how the registration policy ciphers and all this kind of carrying on was gonna happen. And so we've taken a bunch of that stuff out of scope. And so all the security context and everything will be done beforehand. And what we did is we actually did this using a RADIUS EAP registration flow, and that is all in there. Certainly, you can use something else. It's just an illustration of how a registration can be done with a relatively well understood registration model such as radius. Okay? And so as I say, a number of the problems had to do with cryptographic mechanisms, how all that was to be established, and so we have taken that out of scope. Yep. And so, yeah, so we concentrate just on what PDMV two will do. So kind of the the way things go is that you will register your devices. You will decide who's gonna talk to whom, policy, etcetera. And then the measurement, and that is what we are concerned with in the PDMv2 draft is exactly what will be measured, what will be the data structure, and packet contents. And then analysis is also out of scope. Whoever is going to analyze said packets, whether it is in real time or after the fact, be a registered participant, however that policy is done for your network. So and there were a number of key use and layering concerns. So since PDMV two operates at the IPV six layer, we can't simply use TLS or DTLS or anything like that because we it applies across the other upper layer protocols and none of these lower level protocols or upper level protocols are nonunique, and and so we don't reuse any of those keys. The rep the registration exchange is an entirely separate set of events, and that can use whatever mechanism you would like to use. PDMV two is separate from PDM. PDM RFC eight fifty was the unencrypted PDM. And, again, the reason we wanted to have it encrypted is we feel that the data is divulges too much about the nature of the endpoint and can potentially be used for attacks. And so we will ask for a new, allocation from IANA for this, for PDMV two. That needs to be updated in the draft. Yeah. And so those are the two things that need to be updated in the draft asking Ayanna for a new option type and the wire format. K? And so now we are ready for, another review. And if we can get an early review from, I believe, sec area, and, I have to look and see what other area already had comments. But for the chairs, we are ready for that. And if anyone has any comments, please let me know. There are I believe there are implementations of PDM in India already. They used it for measuring DNS performance. We ourselves have a an implementation as well, and I imagine I imagine they will do PDMV two also. Comments, thoughts? I think that's it.

[00:15:25] Marcus Ihlar: Alright. Thank you very much. Any comments from the room? Yeah. Was in the

[00:15:35] Unidentified Participant: yeah. So so first of

[00:15:37] Warren Kumari: all, thank thank you, Nalini, for reviving this work and then we say trying to push it forward. I think that's encouraging to see that the the update and also the the approach you have, I would say, added this so far to to a flow that we say the the contentious part of it. So let's hope that this will make it this time. So, yeah, be before I would say starting again the process here in the work you're working to have the work in plus and so on, make sure that we we have all the pieces we need from the secure there. At least we request security direct rate, Internet direct rate, and ops at least for those. Once we are, I would say, confident that the, yeah, the the new design, I would say, this case is really, yeah, sound, then I will be more than happy to take it back to the AG for review and so on on this. Thank you, again, for

[00:16:21] Nalini Elkins: Yeah. Thank you so much.

[00:16:22] Marcus Ihlar: Yeah. Thank you, Matt. That's exactly what we were gonna suggest here, to to to do a number of early review requests first, and then, yeah, take it from there. Perfect.

[00:16:30] Greg White: Thank you

[00:16:30] Marcus Ihlar: very much.

[00:16:31] Nalini Elkins: Thank you, guys.

[00:16:39] Marcus Ihlar: Next up is, Shailesh.

[00:17:19] Shailesh: Yeah. Good morning, everyone. My name is Shailesh. I'm from Nokia, and I'm gonna present the latest updates, the draft on the characterization and benchmarking methodology of power in networking devices on behalf of all the others. So for folks who are new to this draft, just wanted to give a quick context and recap. So the whole the goal of this work is to, you know, kind of come up with a standardized benchmarking methodology to evaluate energy efficiency of networking devices. Right? So the idea is to, you know, provide a consistent way to characterize power consumption of network equipment under repeatable and well defined test conditions. Right? So the focus is on the device level power consumption under controlled traffic conditions, which enables us to kind of enables two key use cases, which is comparison of energy efficiency across devices and also assessment of the system power consumption and also the contribution of the subsystem components to the overall power consumption. Right? So this is an adopted draft, and it has been presented since ITF one one nine Brisbane and has evolved based on the working group feedback and comments. So we have ever since ITF-one twenty five Shenzhen presentation, we have, you know, revised the version, and one of the updates is that, you know, we have an update on the benchmarking environment and measurement conditions. So we realized that, you know, the power measurement can get influenced by laboratory environment conditions, especially temperature. So which was not explicitly defined, right, in the draft. So now we now the draft recommends laboratory environment conditions based on that CES two zero three one three six. Specifically, the temperature between 23 to 27 degree Celsius, rate of humidity between 25 to 75%, and atmospheric pressure between eight one two to one zero six zero hectopascals. Right? So this kind of improves the measurement repressivity and especially the cross lab compatibility. So the second update we had is over some clarifications over the existing benchmarking scenarios. So for base part procedure, we clarified that all the boards and components must be fully booted before the power measurement is taken. We also added a rationale behind the one packet per second traffic used in the idle plus power. So the, you know, the intent is to send base minimum traffic that is required to bring the networking device out of the low power or the sleep state. Right? So that improves consistency and once again, more reproducibility of benchmarking procedure. So we added a new benchmarking test called typical part, which is basically drawn by the data under representative operation conditions, and that complements the existing benchmark test as well. Right? So we had a couple of more updates to the draft where we strengthen the iMX traffic as the traffic repeatedly repeatability specification. And we also clarified that the non drop rate or the zero packet loss is the default benchmark assumption, and any non zero packet loss must be explicitly declared during the benchmarking procedure. And we also removed the inter packet delay as a mandatory traffic parameter because, you know, iMix is already a well defined traffic profile. So, you know, we thought it was unnecessary and it simplified the benchmarking methodology for us. Yeah. So on the latest updates, we wanted to request the feedback from the working group. And, also, we want to investigate the alignment opportunities with ongoing green working group, right, work. And, also, I wanted to request the working group and also the chairs that if there is consensus, we want to assess the readiness for the working group last call. Thank you.

[00:21:53] Giuseppe Fioccola: Comments on this draft? Yeah. As we mentioned, as we also highlight in the slide, maybe before moving towards the working group plus call, we need more inputs from expert on energy efficiency. Sure.

[00:22:12] Luis: Yeah.

[00:22:13] Giuseppe Fioccola: Because it's it's and it's an expertise that we lock in the NWG IPPM and yeah.

[00:22:20] Nalini Elkins: Right. Sure.

[00:22:21] Giuseppe Fioccola: This is maybe some reviews

[00:22:23] Thomas Graf: Yeah. Yeah.

[00:22:24] Giuseppe Fioccola: Can be useful. Yeah. Correct.

[00:22:26] Shailesh: Maybe in the audience, I might request some experts from Green Working Group to kind of review the draft, maybe made a special request to you to provide some comments in view of the Green Working Group work. Right? So that would be really helpful for the draft. Thank you.

[00:22:47] Marcus Ihlar: Thank you. And next up is.

[00:23:19] Linhui Sun: Good morning, everyone. I'm from lab. I will present updates on their benchmarking methodology for intra domain and the intra domain source address validation. On behalf of my co authors, and the. So let's first overview of the document and have a look at their documents' status. So this document provides benchmarking methodology for intra domain and the inter-domain SAV, which evaluates their style accuracy in terms false positive and false negative rates, control plane and data plane performance, control plane, and data plane performance, and the resource utilization. It adopts a black box and a mechanism and a implementation agnostic approach. This document was adopted by BMW in February this year. And after that, we updated the version zero and presented it in the IETf-one to five workgroup meeting. Based on their relevant discussions in the SAVNET working group, we had updated two versions. So for this meeting, we will present it as a war present to the version level three. This is a summary of the main updates. In the document, we had clarified the scope of intra domain and the inter domain. So based on the discussions about the distinction between intra domain and inter domain style in some network group. And we also updated the terminologies to align with the relevant documents in some network group. And we refined the test methodology and the definitions of their key performance indicators and refined their test cases for intra domain and the inter domain style as well as expanding the reporting format section. This is a conclusion of the discussion on their distinction of intra domain and the inter domain style according to their discussions or their work group last call for their problem statement drafts in some network group. So for the intro domain style, their deployment points is their external interfaces connecting to customer network with no AS and a single host or a set of hosts. For interdomain style, the relearned deployment points is their interfaces connecting to neighboring AS, which may use a public AS number or or a private AS number. So for this document, we use the same concepts to define the benchmarking test cases. We have added some new terminology terms and refined some other definitions of some existing ones. We had add added limited propagation or prefix abbreviated as RPP, which represents the pre prefix configured with no AS path, no enterprise, or other selective export policies. Hidden prefix represents a prefix that is not visible to the routing information, which usually called by their direct Direct Server Return, DSR. DSR uses any cast service addresses while delivering data from edge servers that do not announce these addresses in BGP. We have refined their definitions of SAV related and SAV specific information. SAV related information means the information that is originally proposed for non SAR purpose, but may be used for SAR. SAR specific information represents is specifically proposed for source address validation. For intro intro domain style test cases, we had to remove the test cases for Internet facing and the aggregation router facing network as well removing the test cases under policy based routing and the faster routers scenarios. This align with the definition, so which works at interfaces connecting to a neighbor AS a neighbor network with a no AS or hosts. For inter domain style test cases, we refined their test case for customer interfaces under the scenario of hidden prefix caused by Direct Server Return. We have refined the re reporting format section. There are more test configuration parameters included to improve their clarity. They include the DOT deployment type, intra domain interface type, inter domain relationship type, routing configuration, subtable size, and update characteristics, number of repeated runs, and the 30 just called treatment of the results, some mechanism and the configuration, the test traffic attributes. Okay. Many thanks to all the people for their valuable comments and the reviews on this document. For the next step, we are going to improve the clarity and the completeness of the test test procedures in the document, and we will continue discussions in the relevant work groups and revise the document based on their community feedback. So thank you for listening. I'm happy to take your questions. Any comments?

[00:29:58] Lan Chang: Lan Chang from lab. Just a second, The coverage version of this document is aligned with two important documents in sub networking group. The first one is the intro domain style problem statement, which is already in the IFC queue. And the second one is the intro problem statement document. It's just passed the working group last call and will be submitted to IESG evaluation. So I think the test cases in this version should be stable because the problem is identified. So but we still welcome feedback. So I I think this is a very good chance because the another document is still a value evaluation. So I think there is a good chance for others to look at the new version, and the feedback are welcome. Thank you. Thanks.

[00:30:54] Giuseppe Fioccola: Alright.

[00:31:00] Marcus Ihlar: Next up is Rakesh.

[00:31:31] Rakesh Gandhi: Good morning, everyone. My name is Rakesh Gandhi from Cisco Systems. I'm presenting this MPLS working group document on MPLS IOM on behalf of my coauthors. So we requested the working group last call in MPLS, and chairs asked us to present this here because I am IOM work was done in this working group. So agenda is look at the requirements and scope, the IOAM encapsulation in MPLS, and the next steps. So requirements basically to support the IOM, both the the IOM as well as the DEX option types in MPLS. There are RFCs nine one nine seven and nine three two six that define those option types and using the MNA in-stack and post-stack network accents. So the IOAM is realized as post-stack MNA, and the DEX is realized as either in-stack or post-stack MNA. And the incremental trace option is not supported. There's a typo in the slide. So the first one, IOM and DEX in post-stack, there is a in-stack MNA that contains the opcode that there is a post-stack data for IOM, as well as where the data is, the offset, as well as the scope of processing, edge to edge or hop by hop or select, and what to do if it's unknown on the node. And the actual option type that's defined in nine one nine seven or nine three two six is carried in the post-stack. And it's just an encapsulation there with a post-stack base header, but the option type is carried as is. And DEX is also supported in ISD. Basically, the same option type is carried in in-stack, but a bit of a formatting change because of the s bit and the b s this BSPL aliasing. But there is a flow ID as well as sequence number carried the same way it would be for other data planes. So the MNA has scope for I2E, which is end to end or hop by hop skill select. And based on the scope, appropriate ops and type is carried in the MPLS packet. So requesting a couple of network action op codes for the first solution as well as for the DEX in ISD. So we welcome your review comments and suggestions. IOM work was done here. So I'm sure there are a lot of good ideas in this working group. And we have requested a MPLS working group last call in the MPLS working group. Thank you.

[00:34:48] Thomas Graf: Thanks a lot, Rakesh. I think this is very important work. It brings IOM into MPLS, and I understood you're aiming for working group last call. I think for the IPPM and BMWG community, this would be the right time now to review the document and to give feedback. Thanks.

[00:35:10] Rakesh Gandhi: Thank you.

[00:35:15] Marcus Ihlar: Alright. Next up is Greg.

[00:35:26] Greg White: Good morning. Greg White from. So a brief update on the stamp class of service extension for ECN. This is a draft that was adopted by the working group last fall, I believe. And the purpose of the draft is to define or yeah. Define a stamp extension that allows stamp to be used to test traversal of ECN in both the forward direction and the reverse direction. There was originally a class of service extension in stamp that unfortunately only allowed the forward path to be tested with ECN. So this draft revises the class of service extension to handle the reverse path. So the current status, I posted a draft zero one on Monday. Only two changes. The first one is a technical change. There was an ambiguity about what a session reflector should do if it for some reason, is unable to set the ECN value that was requested by the session sender. So to make that deterministic, we're mandating that zero zero value or not ECT be used in a reflected packet. And the second change is to add a reference to a TSD working group draft that is nearing publication. I think it's in the editor queue that talks about how to use or how to set the ECN value in different operating systems. In addition, there's been some discussion on the mailing list. Actually, it was during the adoption discussion for this draft. An alternative encoding of this class of service information was suggested. We haven't reached consensus on any change there. Through the discussion, it's essentially the the main difference between what we have in the draft and this alternative proposal is the alternative explicitly includes the diff serve code point and ECN values that are being used in the test in both the forward path and the reverse. And that would enable an on path observer to understand the the test and be able to detect manipulations of the DiffServ Code Point where you see n value different stages along the path, which is an interesting feature, for sure. The main problem with it is it throws away the backward compatibility. So with the current draft, we have backward compatibility. A session sender can talk to a session reflector even if one of them supported the old version of of the class of service TLV and the other one supports the new version. So I expect continued discussion on that on the mailing list. Hopefully, we can reach consensus on that and move on from there. There's also some off list discussion between the authors, and Will Hawkins had raised an issue where and this actually was an issue with the original class of service TOB as well, that the behavior was undefined if the session reflector can't use the requested diff serve code point or the received diff serve code point for its reflected packet. And we have a little bit of analysis of that on the next slide. Also hoping to do an interop at the next hackathon. So if anyone has a stamp on implementation and would like to work with us on interop testing, please join us at the next hackathon. So this is that one issue that I mentioned. So currently, in receive stamp packet at the session reflector, there's the DiffServ Code Point on the IP header, and the session reflector inserts that as value referred to as DSCP two in the reflected CoS TLV. And that the received class of service TOB also contains a DSCP one, which is the session sender's requested value that the the reflector should use in its IP header. And if it's not able to do that due to limitations in the the system, the session reflector system, or configuration that prevents use of that DiffServ Code Point, The draft currently, and actually the original class of service, RFC, says to use the received IP DiffServ Code Point. And it sort of leaves it ambiguous as to what to do if the session reflector can't use that DSCP two value. So listed out four options, and probably need to take this to the mailing list to to kind of do the analysis of which option is best. I think we're sort of leaning towards option two or three based on just discussion between the authors and and and Will Hawkins. But be able to get to the list. One other point is that that that class will serve as v two proposal. It doesn't have any ambiguity here. The session reflector can just list what code point it it did use in its response. So a session sender can get that information and know unambiguously what the the session reflector had used. So that is the status. I can take questions if there are any.

[00:41:39] Marcus Ihlar: Any questions from the room or from remote?

[00:41:48] Wei Chang Sun: Looks like no.

[00:41:49] Marcus Ihlar: Alright. Thanks a lot. Oh, wait. There. No.

[00:41:53] Rakesh Gandhi: No.

[00:41:53] Giuseppe Fioccola: I think it's quite interesting. I I just want to point out a draft about the the combination of active tools and hybrid measurement to also to make hop by op measurement to combine. So maybe this can can be mentioned as a possible improvement. So you have you can also make compact measurement and hop by op measurement by by leverage the hybrid tools in combination with stem. So that's quite interesting also for the l four s application.

[00:42:32] Greg White: Right. Yeah. So that that is definitely an advantage of that cost v two proposal. So I agree. I I think that would be a nice feature. So

[00:42:42] Giuseppe Fioccola: I will point out the draft, so maybe you can extend in one new section to to mention this.

[00:42:50] Greg White: Okay. Great. Thank you.

[00:42:58] Marcus Ihlar: Okay. Next up is Bin Young. So we'll talk about a suite of documents, that was originally post for CCAMP, and now it's a bit of a split.

[00:43:16] Bin Young: Good morning. My name is Bin Young Yeon from ATLIST. I will present the three document on the collections measurement on behalf of The Us and the supporters. And the history. This work started a single document, a young data model of PM streaming at CCAMP two years ago. And just to be put working adoption poll conclude, I got a comment. It contained the genetic performance management requirement out of scope for the CCAMP. Maybe it's the IPPM working groups, maybe responsible port. And so we had online meeting and the email discussion with the two working group chairs and ADs. And so finally, it was agreed that a generic part should move to the IPPM working group and the wider technology specific and transport networks remain as CCAMP. Now we propose the split into three document. And the two document for this IPPM, one document for the CCAMP. And especially is the I like to thank to the chair and ADs and to the including Chinos. Give me some opportunity to start sharing my work with you. Especially thanks to Thomas because he's give me a clear guidance which working group I should go with this draft, including many things.

[00:45:09] Warren Kumari: Okay.

[00:45:12] Bin Young: Original document is mixed to genetics young data models, our PM collection and the PM interval capability. And then technology specific young model of PM parameter and the tech transport specific. And so the the two genetic models was taken out to move to the IPPM working group as a separate document. And PM's parameter technology space young data model, remained the CCAMP. And the outline of the three document. First, the document and the collection measurement And for the IPPM standard track, it include the genetic core models, including subparameter, sampling collection, interval hierarchy, three types of collection measurement. And the second document about the interval capabilities. It augment the first young data model and make the server advertise the interval capability to a server. The third one's PM, streamings on transport and the folder IPPM, SCAT document. It import to the first and the second young data model to specify this how to support PM streamings in transport equipment. It also define this the command transport parameter modules and the profile. K. The first document in detail, it described the three stage of pipeline of performance management processing, reading the equipment, and the sampling is the first state. And this model this draft model, the second collections stage. And it processed raw performance date, captured at the first stages, and produced structured performance data and get removing the the unnecessary performance data. And then performance data is delivered to the client. And this draft also defined three collection types counts as captured the total number of event accumulate of a collection interval. Snapshot and captured instantaneous value at uniform times every interval. Third trademark is the captured minimum, maximum value every inter burst with the threshold. Okay? This draft is the proposed hierarchical structure, young data model, and starting from the profiles and the parameter and intervals. And that right figure shows some hierarchical relationships between the sampling interval and the collection intervals. And the, for example, our second and network operation center can be monitored by the counter measurement. And we did a one second interval over three different collection intervals for the d different purposes. For example, one minute's collections for rapid default allot, second fifteen minutes for the maintenance. And they show the some relationships that one sampling interval that can be mapped into the multiple collection interval as a hierarchical architecture. And that this slide those slide, they show the relationship between the term terminologies used in draft and the terms in the IPPM framework working document. And so the terminologies of collection types should be aligned with the ITAP IPPM store framework document. And for example, snapshot correspond to singleton and counts and the title marks correspond to statistics of a sum and the min max as a temporal aggregation and r f c sixty nine ten. Also, did this this draft is to try to minimize collection type, strict collection types, and the base on the G.7710 because it picks the small operational valid dataset instead of all possible statistics, including average variance at the statistics and not include this in draft. And it so it keeps any side, the computation light, heavy analysis in the external system. Okay. Second document is about advertising the interval, the service supported. It augmented ITEB system capability, IPC 91, 96, and so that it also mirrored the first young data models of profile and parameter structure. There are several regions why the discovery of interval capability is need Because one parameters that can be measured by the multiples collection interval. Sometimes sampling interval can be variable. In the in this case, server should know the capability, interval capability server can support, you know, advanced people computer configuring the collection measurement. Okay. This draft defines the metrics of the interval capability interval capabilities with the mean max value and unit and the granularities. Okay. So the document is about applying the first two model to transport the equipment by importing the first and the second young modules. It also start to find the three group of performance parameters. First groups can be the monitored by one of the three collection types with a fifteen minutes for the maintenance. It can be the short term operations, rapid fault detection. Second groups can be monitored by the one of the three collection types. We did a twenty four hours collections for the maintenance. For example, daily trend, strategy planning. And the third ones can be monitored by the only count measurement with the twenty four hours for the QS monitoring. That can be the SRA compliance and the QS reporting. And this draft also addressed the whole press procedures of end tender streamings from the discovery interval capability and streaming the performance data. It start the it start at the first stages, the end of the and the servers advertise the interval capability toward the servers client, and the client, and based on the information, subscribe the performance data and the interest in. And then after configured, and the servers stream the performance data periodically. After that, and some time, threshold report is triggered. Stand when the threshold values across the the monitoring value cross the threshold value. Given the nearly two years of review in the CCAMPs, we we believe that the work is mature. We do welcome to prompt working adoption call. The reason is that I'm a little worried about another two years IPPM reviews. And, anyhow, I want you have a deep comment, a deep consideration on this document, and all commenters always welcome. Be putting the GitHub links below. Thank you.

[00:54:52] Thomas Graf: Thanks a lot, Bin. And absolutely right, I think there is a lot of work already in those documents. And I think we will make sure on our side that we can quickly move forward to an adoption call. I have one question. I was checking on whether there are existing ITOT liaison statements on this work, but I haven't seen any. Are you do you know, are there any liaison statements currently open? If not, I will clarify with the CCAM chairs.

[00:55:26] Bin Young: Not yet. And so I'll try to see, you Okay.

[00:55:29] Lan Chang: Good then.

[00:55:30] Thomas Graf: Good. Okay then. I think we're gonna initiate something just to be sure that the ITO community is aware of of this work.

[00:55:37] Bin Young: And so the revisions that is planned to start by next year. Now, currently, we are working on the G.7710 document test.

[00:55:47] Thomas Graf: Okay. Perfect. Thank you.

[00:55:52] Giuseppe Fioccola: Okay. There are other questions. One is from me. So I'm wondering whether during the processing of measurement, did you consider how to categorize the data eventually differently depending on whether they are performed with active, hybrid, passive methods considering the IPPM terminologies?

[00:56:17] Bin Young: Yeah. So this draft test does not consider the any OAM mechanics in band, out band OAM. Just, you know, after getting this OAM data or performance data?

[00:56:30] Giuseppe Fioccola: No. I mean, I mean, the data, classify the data maybe with the tag if they are performed with different active because the information can be useful when you collect based on the kind of methods.

[00:56:45] Bin Young: Some metadata is Yeah. Necessary?

[00:56:49] Giuseppe Fioccola: Probably. Yeah. We can further discuss Yeah. Because to align also with the PPM

[00:56:55] Lan Chang: Yeah.

[00:56:55] Giuseppe Fioccola: It terms and maybe it's an information that can be added as metadata. Yeah. Yeah. Yeah. So that's something to consider. We can

[00:57:04] Bin Young: Yeah. Yeah.

[00:57:05] Giuseppe Fioccola: Because I didn't look I I I will look more deeply to the to the end because from a quick look, I didn't see this categorization. Maybe

[00:57:16] Bin Young: it's useful. Yeah. If you look at the my suggestion, there might be the profiling of the set of performance parameter. Yeah. In that case, it's some has some characteristic of the performance.

[00:57:30] Giuseppe Fioccola: Yeah. Something like that. Yeah. Yeah. Okay. Well, maybe some explanation would be good then, some additional information on that.

[00:57:38] Chin: Chin. Yeah. Chin, actually, thanks for this update. Actually, I break it down into three draft. We saw a lot of overlapping issue. And my comments, you know, for the you know, you have two draft or, you know, both actually is a young data model. One of the young data model is a system capability. I'm running wondering whether this should be, you know, the dev developer in the NetComm or this should be combined into the single draft, you know, just, you know, feel the chair need to decision make a decision on this. The second is there's a error map or working group. They already get conclude. They define the error map or main and the young data model. And I see a lot of, you know, maybe some building block you can reuse so you can check whether some there's something overlapping or you you need to clarify how this is related to LMAP or young data model.

[00:58:37] Bin Young: Yes. Okay. Let me know the the detail. And anyhow, the the first document is the collection measurement. Mainly, focused on the collection measurement based on the G.7710. Second one is really different. Right. Capability issue. Not this come based on the GDA seven seven ten. It's a it's a little bit separate issue. Okay?

[00:59:03] Wei Chang Sun: Okay. Thank you. Yeah.

[00:59:13] Marcus Ihlar: Alright. And next up is, launching team.

[01:00:04] Lan Chang: No. No. This one. The benchmarking methodology for relying party.

[01:00:10] Giuseppe Fioccola: No. RPK. RPK.

[01:00:12] Lan Chang: RPK relying party.

[01:00:13] Giuseppe Fioccola: Yes. Money presentation.

[01:00:16] Marcus Ihlar: It's too small window

[01:00:17] Rakesh Gandhi: here. Yeah.

[01:00:30] Lan Chang: Hello, everyone. I'm from lab. So this draft, describes the benchmarking methodology for RPKI relying party. So here is a brief introduction to RPKI and the relying party. RPKI, resource public key infrastructure, is a public key infrastructure that represents the allocation hierarchy of IP address space and AS numbers. And then resource holders can use API to publish signed objects such as to provide the authorized information for BGP route filtering. And relying party, the RP, links the RPKI repository to routing system. It's a very important role. It retrieves RP objects from the distributed RPKI rep repos and validate these objects and generate the, validated outputs such as the validated raw payload for use by routing systems. So why do we need relying party benchmarking? There is a answer. Currently, there are various relying party implementations, but they may differ in their internal validation workflow and processing strategies. And they may differ from some RPKI, IFC requirements. And we note that some deviations from IFC requirements can be difficult to identify if we don't have a controlled repository or without well defined test cases. And the benchmarking methodology is also useful for both relying party developer and a relying party user. For relying party developer, yet for me, like me, I'm a human, not a machine, so I may miss or misunderstand some requirements, standards that in the RFC. So I may do some some something wrong in my implementation, but I didn't know that if I don't don't have a benchmarking methodology to test it. And for users, for operators, they also want to know, how do I select one that meet my requirements among all of those, available RPA implementations? How can I just select select to to choose the one I like? So this draft answers the question. How can we consistently evaluate and compare the correctness and the processing performance of different real time party implementations? So this document includes the benchmarking methodology in a controlled laboratory environments, and it's a black box box test based on only external internally observable outputs such as the VIP sets. And after before we write this document, we share our thoughts with several RPKI relying porting teams, and the feedback is that the correctness is the most important property. And the performance improvements are meaningful only when the relying party produces the correct validation output. So we consider functional correctness for required validation procedures as a primary evaluation objective, and we also measure the performance. Here shows the black box benchmarking test environment. Here is a controller and servers and software under test, which is also a reliability. The servers host and publish the controlled API repo, and the the software under test retrieves and validates the objects in the servers and provide the validation output to the controller. And for the controller, it controls the repository content and start tests and compare the run party's output to raise the expected outputs. The general test methodology is very straightforward. We first start from a no invalid repository generated by the controller, and we run a test. And then we can got the valid VRP set. And after that, we will modify one object or modify a test condition at one time. And, for example, we can introduce a malformed raw or malformed manifest. And then we run we we perform a code start run again, and then we can compare the generated VRPs with the expected results. And all of our functional correctness test cases are derived from validation requirements defined in the relevant RFC, and we try to emulate whether a relying party correctly performs the required validation procedures and produces RFC conformant outputs. And for for example, we can we we will test the d r encoding correctness, the syntax validation correctness, and the certification pass validation correctness, as well the signed object signature validation correctness, and also for some processing correctness or manifest to CLL or other types of RPK objects. And for each test case, we define the objective, the procedures, and our expected results. Due to time limits, I cannot introduce all of the test cases, but here are two examples. For, in the first example, we first built a valid baseline repository and run the down party. We can go to the valid v VIP sets. And after that, we can modify one lower, make it malformed, such as we introduce a wrong, syntax format to use a wrong signature or do some issues to bad issues to that API object. And then we run we perform a code start to run again, and then the expected results should be that the VIP derived from that malformed ROAS should be absent, but other VIPs should remain. And we also test the performing time for run on porting, but in current version, we only focus on one short validation rule. And we, and the controller will calculate the time from when it starts the run to when the SUT provides its VIP set to the controller. And the controller can decide how the VIP set is fully received. So the in next step, we plan to expand the test coverage by adding additional test cases, such as the CRL processing, correctness validation, and the processing time test and the daemon mode operations. And we also welcome feedback, and there are two questions for the community to consideration. Yeah. And we thank to and the Fort RP development team and the for their reviews and comments. Thank you.

[01:08:48] Giuseppe Fioccola: Thank you. Question, comments? Yeah. If no one I have I have a question. I just noted that you also present inside the ops this week. Right?

[01:09:07] Lan Chang: I I I I requested a presentation, but that didn't

[01:09:12] Giuseppe Fioccola: There was no time?

[01:09:13] Lan Chang: Yeah. It's not included in the agenda inside the office.

[01:09:16] Giuseppe Fioccola: Ah, okay. Because yeah. I was curious. I was wondering what was the feedback from side the ops because since the experts are there.

[01:09:24] Lan Chang: That's what I sent email on the side of main

[01:09:26] Giuseppe Fioccola: list.

[01:09:26] Shailesh: Okay.

[01:09:27] Lan Chang: Yeah. And I I see Misha here. So yeah.

[01:09:30] Giuseppe Fioccola: Okay.

[01:09:33] Unidentified Participant: Okay. Thank you.

[01:09:35] Marcus Ihlar: Alright. And then we welcome Rakesh up again. We can talk about stamp MPLS header. So this was originally part of a common document that was split and that one part is through working group last call, and this is the other part.

[01:09:52] Rakesh Gandhi: Yeah. Good morning, everyone. My name is Rakesh Gandhi from Cisco Systems. Presenting this TAM for MPLS extension headers on behalf of my co authors. So we look at the requirements, updates, summary, and next steps. So as Marcus mentioned, this was part of a working group document, and we split it into IPv6 and PLS about a year ago. And IPv6 just completed working of last call and MPLS. I think it's in the state where we made the parity with IPv6. So they both should be in the in the exact same state except different data plane.

[01:10:43] Giuseppe Fioccola: Do you want to merge again? I'm joking.

[01:10:46] Rakesh Gandhi: Yeah. It's in the IESG now, so it's a good idea. So requirements are basically using stem for hop by hop and s two s measurements using m and a. So in in some way, this is the use case document for the MPLS IOAM solution that I just presented few minutes ago. This one shows how we can use that MPLS IOAM with stamp to do hop by hop and s two ways measurements for both one way and two way. The scope is really stamp and all of the m and a as well as IOAM RFCs. So just to give a summary, there is m and a in stack. There is m and a post stack that carry various IOM information. And the payload is stamped in this case, and there are stamped TLVs. We do hop by hop measurements. And on the reflector side, everything that's collected in m and a is copied into the stamp and send it back to the sender. There's a header control. So if you want to do measurement in the reverse direction, we also add the MPLS headers in the reverse direction in step. So that's the gist of the solution. And I p v six works the same way for I p v six extension headers. So we made quite a bit of updates to keep up with I p v six. This was presented in Montreal. Since then, Fabian has joined. Actually, Fabian implemented this also in open source and gave very good feedback and improved the draft. So thank thank you, Fabian. We presented this draft also in MPLS working group session to get their feedback. The same way we did for IOM presenting here, we did this one in other working group in. So many updates to keep up with IPv6. There is requested bits that's defined in the header, the clarification of conformance, how the MPLS header is processed on the egress node. We clarified that as well. I think Greg had comments on that. There is a m and a header control sub TLV. It's used to add the headers in the reverse direction and conformance for it. We clarified all that as well. And then operational and security considerations also updated to match the IPv6 part. So we're asking for two code points, one a stamp TLV and a sub TLV for the type 12 in reflected packet control. There is an implementation open source. Code is available. Thanks to Fabienne. So we welcome your review comments and suggestions. This draft was part of our working document. There is an implementation, so I think it's a good time. We we did request the working group adoption on the mailing list. And thank you.

[01:14:08] Marcus Ihlar: Alright. Thank you. Do we have any comments from the room, from remote, from anyone? If not, it's very encouraging to see that there is, open source implementation of this. That is very much appreciated and in line with how we want to see things move forward. So this is definitely one of the priority items for us to adopt. So we'll take that into consideration.

[01:14:37] Thomas Graf: Thank you.

[01:14:38] Marcus Ihlar: And, yeah, you can stay on.

[01:14:42] Rakesh Gandhi: Again, I'm Rakesh Gandhi from Cisco Systems presenting the draft on residual BER using STEM. We love STEM. It's simple. On behalf of my co authors listed here. So agenda is look at the requirements, updates, summary implementations, and next steps. So this is a residual b beta error rate measurement. So that's the beta error that escapes the CRC detection and FEC correction and result in UDP checksum errors, for example. And we measure in both directions, And we want to get a very good approximation of the residual bit errors, and this is a stamp. So we welcome Lee as co author. Looking forward to working with you, Lee. We addressed comments from Ernesto, Carsten, and Lee as well. We added new TLV. There was an ask for it to detect the burst in bit error, so we added that. We clarified various things including unknown flag and the the pattern must be integer. The pairing must be integer of the pattern size and many other editorial changes. We are also added Cisco's shipping implementation. We're shipping using experimental code points at this point. So just to give a summary of the procedure, stamp is used with extra padding TLV, and padding contains the a bit pattern. That's also signal or locally configured, and reflector basically matches the pattern with in the padding and looks for the bit errors and the burst and puts the data into the stamped TLV and sends it back to the sender. So that's the gist of it. There are three TLVs defined for the pattern, the beta count, and the burst. So Cisco does have implementation. It's shipping. There is a CCO documentation on how to configure it and use it. There is also open source implementation. Thanks to William. It's available. All using the experimental code points at this point. So we welcome your your comments and suggestions. We requested early allocation of IONA code points on the mailing list, and we have also requested the working group adoption on the mailing list. And that's what I have so far.

[01:17:18] Richard Foot: I forgot to raise my hand. I'm sorry. Okay. Richard Foot, Nokia. I second your early adoption of a code point. It's gonna make it much easier, coauthor, to interoperate and not have to transition. So if we could really push for that, that would be very helpful.

[01:17:38] Rakesh Gandhi: Thank you,

[01:17:40] Marcus Ihlar: Yes. We can help with that. Lichang, in queue.

[01:17:49] Lichang: Yes. Lichang from Huawei, and thanks for addressing my comments. And I think the draft now is useful and mature enough for adoption. And one more suggestion is maybe we can create a git GitHub or to to collaborate with other callers and contributors. It can use it to record their issues and can check the updates. It is very useful. That's our suggestion. Yeah.

[01:18:26] Marcus Ihlar: Yeah. That that is a very good suggestion. And Yeah. One suggestion we could have if if we adopt this is that we actually have a a repository or we have an organization for for the working group that we will is now called IPPM. Soon will be soon will be renamed, and we are more than happy to to host repositories, if that would make sense. So yeah.

[01:18:49] Rakesh Gandhi: So is this the first one getting adopted as draft, BPM something?

[01:18:55] Marcus Ihlar: That's a very good question. It depends on timing here. We're we're we're, yeah, it will be a little bit of process here before that happens. But yeah. Thanks a lot.

[01:19:05] Rakesh Gandhi: Thank you.

[01:19:17] Linhui Sun: Hello, everyone. From lab. I'm going to present to the updates of their benchmarking methodology for router validation. This is a joint work with from Huawei. So let's all review the document and have a look at the draft status. So this document is to define a repeatable and controlled benchmarking workflow for routers which implement their ROV and provide a standardized metrics that allows a complete the complete reason across their ROV products and the software versions. And it keeps the device and the test at a center and assurilate their noises because we want their observe the differences to reflect our behavior itself. We uploaded the version zero zero before the ITF one two five and presented it during the an ITF one two five work group meeting. So during the meeting, we received some comments from Kristen. And based on the comments, we update tier to the version there one. After their meeting, we also introduced this draft in their set ops mailing list, and they required the experts to help review their document and the seek feedback. So based based on comments on Maria from team, we updated two versions. Today, I will present the version zero three. So let me repeat the motivation for this draft. Our way department is increasing, but they are in no standardized device level benchmarking methodology. So we need common inputs, timers, and metrics to do their comparison between different ROV devices and the software versions. And there's a VIP skill, update churn, and the CPU memory cost. ROV have operational effects, so this should be tested. So this document filled a clear gap. There is no standardized and the repeatable way to benchmarking the RV behavior on routers. We have four key components to build their benchmarking environments, including RTR emulator, BGP generator, DOT, and a data plane tester. The RTR emulator is going to mirror a real cache or generate a synthetic VRPs over, like, left cases and the control to reset. The BJP generator is used to provide a baseline read and injects they inject the selected prefix of whose data will be affected by ROV. The data and tester is used to generate control of forwarding load to study how their traffic load will be affected by the control plane ROV behavior. So we also provide some requirements for the workload. For example, the RTR emulator should support controlled bursts, such as 100 to 10,000 VR peers per second. And the the BGP generator should should use us false should use a a stable full table read and should generate advertise meaning prefix. For the benchmarking tests, we apply the three types, including latency oriented, scalability, and robustness. The latency oriented test consists of ROV update processing latency, ROV validation latency, the BGP convergence with and the result ROV. The the scalability includes VRP scalability and the capability to pause process VRP churn and the resource utilization. For the robustness, it's mainly about the RTR session behavior under different few OR scenarios. So this test answers three questions. So how fast is ROV processing, and the way is a scalability limit, and does our in ROV implementation remains stable during the faults and the chain. We have we still we're here with summary of the comments we still on version zero zero and our revisions based on the comments. Comment month suggest suggest our two consider the RTR disk scale and structure. The parameters include the total count or ROA, the percentage of the whole address space covered by the ROAs compared to the address space announced by the routers. And we are also reminded that a chosen ROA may be covered by another ROA. So we extended as a VRP dataset input conditions. So in the in this document, it asks the test to report not not only the VRP set size and the churn, but also the total number of ROAs, VRP coverage relative to the announced routing space and the VRP overlap characteristics. The second moment is about the section 6.1. It may be different whether one single RTR update should modify one or more VRPs. Also, it's a different fall invalid by bad ASN or invalid by too long. So in the in the current draft, we clarify the VIP update and the validation change characterizing characterization. In section 6.1, it clarifies whether an RTR update affects a single VIP or multiple VIPs. The document also ask the test document to the type of validation, state change, triggered, such as origin AS mismatch or prefix lens violation. Common three, we are suggested to to we we have we have received some suggestion about the measurement data. The measure the data may be different based on the scattering or the change of VIP. So we agree with that. So in section 6.2, we added the affected prefix locality control. So we required the test through the report whether affected prefix are localized within the same region or distributed across different regions. We have received two comments about section six six point three about their test procedures and the timer to calculate the BGP convergence time. We have refined the BGP convergence test in section 6.3 now. It separates VRP triggered BGP convergence from BGP triggered convergence with and without ROV. For the last comment, we suggested to measure the full BGP recalculation after the RTR convergence. We agree with that. We attended the RTR recovery measurements in section six point six point seven. We now include a BGP recalculation and a convergence after RTR session recovery, cache fill over, and the full resynchronization. This is a summary of of the main updates on version zero three. So, for the next step, we see comments and feedback from, BMW, IP APM, and CDOPS, and the collecturations are welcome. And we are going to finish our prototype invitation and the participant in the IPF one two seven hexam. Okay. Thank you. I'm happy to take your questions.

[01:28:46] Marcus Ihlar: Okay. Thank you very much.

[01:28:48] Linhui Sun: Thank you.

[01:28:49] Marcus Ihlar: Next up is.

[01:29:11] Presenter: Good morning. I'm from Shanghai Jiaotong University. Today, I will introduce a proposed switching efficiency framework for assessing how effectively AI design center networks turn its deployed switching resources into useful computational progress. In our data center spotting LM training, both communication algorithms and network architectures can shape and network efficiency. First, consider the already primitive in a ring based implementation. A GPU may receive intermediate or redundant data that is later discarded during reduction. This creates network and endpoint activity that does not directly become useful computational input. With in network computing, reduction can be performed in the network so that GPU received only the final reduced result. This tool implementation can therefore exhibit very different efficiency characteristic. The same issue appears at the network architecture level. A three d Torus can efficiently support communication between neighboring nodes by its distributed symmetric switching capacity may be a poor match for workloads with highly uneven traffic demands. By contrast, a rail optimized architecture with centralized higher hierarchical speech in capacity can better accommodate such demands by may incur additional multi hop forwarding. These differences may not be visible from traffic volume or utilization alone. Our question is, can we define magic to quantify network efficiency for LM training? Fire output and bandwidth utilization are useful starting point. However, they show only how active the network is. For LM training, an efficiency metric must meet further requirements. First, the metric should distinguish computationally effective data from other traffic. Computationally effective data is the data retained and used as input to subsequent neural network computation. Not every part carried by collective communication meet this condition. In ring or reduce, for example, and reduce intermediate data consumes network resources, but it's not retained after the reduction. A metric that treats all carried by equally cannot capture this difference. Second, the metrics should capture workload topology fit. DP, PP, and TP in training produce traffic that is often sparse and uneven. The metrics should indicate how well the deployed network topology and capacity support the actual training traffic rather than relying on static bandwidth properties. Third, the metrics should be diagnostic. Beyond providing an overall efficiency score is to help analyze sources of inefficiency and inform optimization efforts. These requirements motivate the proposed switching efficiency framework. Switching efficiency, eta, is the computationally effective data in equipment per unit time divided by total switching capacity. The left side defines the numerator. D is the full gradient or activation tensor by GPO, and delta d I is the data equipment from communication primitive I that is retained for subsequent computation. Its value differs across communication primitives. Intermediate or unreduced data may consume narrow resources, but does not contribute to the numerator unless it is retained. The right side defies that numerator. Some of RP as the theoretical egress rate of the included switching resources from SPI and LIF switches to in server switching. Eta therefore measure how effectively the deployed switching resources turn into effective data fire output. Eta extend decomposed into three factors, which show where the efficiency loss occurs. To calculate these factors, we measure four quantities, c total times t, the total theoretical egress byte over the observation window. We forward egress bytes emitted by the measured falling port in the measurement domain. We receive by successfully received at the measured endpoint, and we effective retain data used as input to subsequent computation. The ratio between these four quantities used the three factors. Post utilization as how much aggregate capacity was used. Routing efficiency as how much forwarding work was required for data received at endpoint. Data efficiency as how much received data was retained as effective data. A low port utilization point to idle or poorly distributed switching capacity. A low routing efficiency can reflect the tours with transmissions, loops, or topology traffic mismatch. A low data efficiency prior to redundant or unreduced data. This factor separate capacity use, routing, and data efficiency causes of a low result. This NSF-three all reduce result show how the metric set exposes the source of an efficiency gap. We evaluate five all reduce algorithms across 16 to 128 GPUs on eight GPU NVSwitch servers connected by a single inter server switch. In this evaluation, Etaz follows exactly the same trend as algorithm bandwidth, a commonly used communication performance metric. Eta, therefore, preserve the familiar performance view while quantifying each algorithm's remaining gap to the network's ideal or reduced performance. The decomposition as the efficiency gap exploration, Theta and gamma show where where whether a low eta is associated with unused switching capacity or with received data that does not become effective. This reveals why some algorithm achieve higher or lower algorithm bandwidth than others. Four score experiment makes this concrete. It's fatter. It's lower than that of most evaluated algorithms and declines as GPU scales grows, revealing underused switching capacity. Its gamma is also substantially lower than MVOS indicating data efficiency limitations. Here, we we evaluate dense and MOE training at 4,096 GPUs under multiple parallelism configurations. Each curve showed the meanwhile across door configurations as one network parameter changes. For dense training on three d Torus, adjusting the bandwidth allocation ratio across different Torus dimensions raises aggregate data from 0.32 to 0.57, indicating better use of provision sufficient capacity. For MOE on rail automized, increasing server size raises both beta and delta. This indicates potential gains from reducing inter server and multilevel forwarding. In network computing affects different factors. For dense model training, the reduced endpoint data that is not retained as computationally effective data, raising aggregate gamma from 0.64 to 0.99. Taken together, this curve show how the decomposed metric attribute the effect of each design change to a specific efficiency factor. This slide shows the same experimental workload inputs as the previous slide and each part representing one parallel some configuration. Data efficiency shows that redundant data reception is an important communication efficiency bollocks for both the network architectures and workloads. Looting efficiency highlights the looting cost of MOE auto auto traffic. It remains low in both error rate architectures indicating high forwarding work provide deliver to endpoints. Port utilization adds a third wheel. Real Automate has higher factor in this evaluation, yet MOE remains challenging for both architectures. To summarize, switching efficiency provides a unified way to quantify overall network efficiency, diagnose and look causes of inefficiency, and guide the design of more efficient AI decentered networks. We would like to invite discussion on two questions. One, how to decide representative workflow to make switching efficiency results more meaningful? And two, how can the required measurements be collected conveniently and reproducibly? Thank you. That's all. Any questions or comments?

[01:39:10] Karsten: Yes. This is Karsten from ENTC. Thank you for reading the statement to us. It looks to me more like an academic paper than an IGF draft. So that's the one thing. So I'm not sure unless the the style of this work is dramatically changed. I think it's not the right place here. And, also, it's hard for me personally to figure out how much of it is actually human work because even you're just reading it, although your English is is quite perfect from the pronunciation. So I think it would be a good rule for the future to require, like, a personal presentation so that we can understand how how much do you actually gather this topic. And I'm asking this because section five in your draft is, depending on which AI detector you ask, between 30 to 100% AI generated. So I'm not sure how much human work is involved in this draft at all. It looks impressive, but for me, the the bad taste remains like maybe just an AI has spilled out the gamma delta and and variables and has come up with these ideas. Can you comment on this?

[01:40:22] Presenter: For the one for the first question, it is we want to proceed this this academic work to to an IETF draft, and we are we are connecting the industry to, like, Huawei to discuss on how to mix this be a more practical metric. And second is I I have I involved the jobs, and I and and now I have to read read this in my paper.

[01:41:13] Nalini Elkins: Okay.

[01:41:16] Wei Chang Sun: Thank you for your comment. I'm Wei Chang Sun. I'm coauthor of the draft. Is my student. So the basic idea of this draft is provide a very high level metric for us to understand how the switching resources in an AIDC is utilized. A very good analogy I borrow from my interactions with an expert during this IETF meeting is PoE. So people, PoE is a very simple yet very communicative metric that people widely use when they design and deploy, evaluate data center data centers. So with, this, switching efficiency, we aim to provide similar metric that probably can be used by, for example, ASIC designers. They are designing ASICs with different relics, right? Small size, bigger size, large port accounts and so on. So when these resources are deployed in air data centers with different architectures, you may wonder how these different architectures, they differ in terms of the utilization of the resource. So we would like to make data centers operate, data centers be leaner and operate in a greener manner, right? In the sense that all the switching resource can be utilized. So this is the basic, the most fundamental idea that we propose. So I would like to say this is the first appearance. We try to bring this idea to the community. Of course, we have to turn this idea, this concept into something that is concrete enough, for example, to make it a software, to implement in testing devices. I had a very good conversation with professor who has a student who is now a leading measurement device manufacturer. So he's very interested in, for example, bringing our idea into something that it can be implemented in a testing device. We have very high capacity device available to test the various elements in data centers. So probably this can be one of the features that can be implemented. Alright.

[01:44:04] Marcus Ihlar: Thank you for, this discussion and for taking this work here. I it's very much appreciated. And please keep engaging in the mailing list, especially if you want to transition this from something academic to something more, you know, ITF engineering like. Yeah.

[01:44:23] Wei Chang Sun: Sure. Sure. Thank you.

[01:44:24] Marcus Ihlar: Encourage you to engage over the mailing list. Thank you very much. Right. And now we're moving into and and we're actually good. We're a little bit on the agenda, but now we're moving into this section of more lightning style talks. For everyone who presents now, please be brief. Two minutes. Do an overview of variety. And yeah. Thank you.

[01:44:53] Unidentified Presenter: Hello. Hello, everyone. I'm from telecom. This document title is f f I've never worked for monitoring pack loss caused by list of congestions. So so so so so I can I kinda count you the slide?

[01:45:31] Bin Young: I

[01:45:35] Marcus Ihlar: Yeah. We'll reshare it. Just a minute.

[01:45:37] Unidentified Presenter: I kinda counted the slide. So next.

[01:45:47] Marcus Ihlar: You should have control of the slides now. Yeah.

[01:45:50] Unidentified Presenter: Okay. Okay. Yeah. Because of I need I need to get time. I only introduced the summaries. This document describe a comprehensive packet loss monitoring framework. The proposals that is capable to determine the time and location of packet loss policies. Accurate as that's the graphs, discarded packets, past what track flows are contained in this kind of pack packets and identify what type flows need to work best, but also after accurate IP loss radio results. More importantly, the purpose process came. Okay. Chill lead or even low interference to network and is applicable to early data player without the monitor modified the waiting chip and packet head. The proposal from from MacMelin compare comprise the network devices and the collection analysis system. The device required to add four new functional functional modules, real time packet loss detection modules, packet loss information reporting modules, cache modules, discarded packets, and particular loss via upload modules. Collection and the analysis system mainly consist five packages. Package loss data collections, measure traffic traffic flow collection, pack packet loss status statistics, and pack loss radio measurement modules. We welcome any comments suggested. Thank you.

[01:47:48] Marcus Ihlar: Thank you very much. Thank you for being on time. Next up is leaving. Yeah.

[01:48:05] Linhui Sun: Okay. One page slide. So this draft, they called considerations for interpreting packet loss observed by active performance measurement. It's going to address the integration problem by measurement by in active performance measurement. So when their active measurement approaches looks at a packet loss, a missing response, or a timeout, the outcome is valid, but the result itself cannot cannot tell us what had happened. So for the same observation, but maybe correspond to different causes, such as congestion, discard, routing, or forwarding, change, endpoint behavior, filtering, or ECMP or link aggregation or validation related that is a card. So this document aim to separate the measurement outcome from the inferred causes and suggest some operational context to help to the measurement analysis, such as the telemetry context, routing forwarding state, counters, or other validation state. This document, the node two define a new metric or protocol. And the question for the IPPM workgroup is that is this guidance useful, and what context should be amplified for the operators? So if we are interested, welcome to discuss it and send your feedback in the mailing list. Thank you.

[01:49:55] Marcus Ihlar: Thank you very much. Next up is Mailington.

[01:50:18] Mailington: Oh, now this last one, I think. Yeah. Okay. From China Mobile. Okay. This is the draft- from scenario that security evolution benchmark for AI agents. We have platform for the agents, and we have showed in the Hexum project. So the goal for us is to establish a standardized benchmark to assess agent security across for critical dimensions. So the framework enables core iterative scoring cross model compare each and the career risk level here. We can see that the four dimension is foundation model, interaction security, and operational security, basic security. These four dimensions and in detail metrics, we have 55, I think. And from the scoring process and waiting, we use metric score and also dimension score and the final score with different waiting here. And, also, we do the two example here is that OpenCrawl. Okay. We can see the version and the. So for OpenCore, the final score is bit row here. So we think the high risk level is really critical from our testing results that we observation that several feature in foundation model alignment and basic security protections and indicating high vulnerability to adversaries attacks. But the HAMMRS, the final score is a little higher than OpenCROW, but it's almost below 60 scores. So it's also the high risk level. So we think the relative between operational security, but significant weakness in interaction, security, like like, a prominent injection defenses. We have individual draft in benchmark working group, and we also have create GitHub repo here for the security evaluation benchmark. If you select the missing mail, you can find the project. And, also, I will zip the Hexum project and also the GitHub address to the benchmark working group. Thanks.

[01:53:23] Marcus Ihlar: Thank you very

[01:53:23] Mailington: much. Really sharp. And

[01:53:28] Marcus Ihlar: next up is Luis.

[01:53:42] Luis: Bye bye. I will just take it there. One one minute. Yeah. So thanks so much. This is Luis from Telefonica. I will introduce briefly this initial idea of how the a benchmarking methodology for a agents in network operations. So, basically, now that we are exploring this this way of introducing agents for doing network operations, so troubleshooting, configuration, checking of the configuration, optimization, yeah, at the end management things. So to to have a a kind of methodology so that we can benchmark the the behavior of the agents and and basically understand if there are drifts, if there are some bias, if there is some performance issue in in a, yeah, multi vendor implementation of these agents for specific tasks. So the the initial idea is to, basically, to introduce the the point. We are defining in the draft trying to to define what could be the system under test. So that would be the a agent with a specific mission together with other tools that could be required for this agent to to perform. And, also, the the the idea will be, okay, consider the the the the model, the version, the the configurations, the prompts that should be used for the specific testing for that specific mission or action that we are validating. Also, we are considering to define the benchmarking framework. So what will be the the definition of the task that will be accomplished by these agents in the in the methodology so that we could define some reproducible way of of testing independently of the implementation of the of the agent. And here, we could things that we are proposing is to explore either the the functionality of the agent as a stand alone for with a specific task or even running multistep workflows or how this agent is capable of defining the workflows to be run so that we can essentially benchmark different solutions provided by all different ways of reasoning of the agent according to the different implementations. We are also proposing a number of of metrics, performance metrics, of course. So time of completion, time to first action, and so. Effectiveness metrics as well. So the the rate of success at the time of accomplishing the task, the degree of pass fail ratio in terms of this task and and so. Efficiency of it, an important thing to to consider is that they they consume tokens in the actions in the all the reasoning of the agent and so. Also, the location of tools, essentially to understand what is the the behavior. The robustness, so how the the agent performs against the any distortion in the in the inputs that they it could have it could receive for performing actions, stability and so, and finally, the coverage. So what are the tasks that is being handled by the agent? And, yeah, basically, how we address the how it address the problem. And finally, we are also trying to define the what will be the rep the format of the reporting of the of the test. And this would include the description of the agent that is the system under test, the benchmarking frameworks because we could have kind of different missions for the agents or the definition of the task, the execution model, the logs, and and the environment. Then the the best marketing metrics that we have referred, and and finally, the time to run and other relevant conditions, like the basically,

[01:57:05] Lan Chang: what

[01:57:05] Luis: is the results, what are all the metrics covered, and so. The reporting that could be maybe following some JSON or YAML format or whatever. So depending on on how we discuss what could be the proper way of of doing so. If we are thinking on automation capabilities, probably YAML could be suitable for it. And that's all for for my side. Thank you.

[01:57:25] Marcus Ihlar: Thank you very much. We see you have one person in the queue.

[01:57:28] Lichang: So if you have

[01:57:28] Marcus Ihlar: a quick question, please go ahead.

[01:57:32] Unidentified Participant: Hey. Hello. I'm from University, and I'm very interested in your talk. And I wonder how can you build benchmarking or test cases that is vendor independent with for the agents. Yes. We find it very difficult to finish this task. Yeah. Yeah. Yeah. So do I have any opinion of it?

[01:57:57] Luis: Yeah. That that would be something that we need to think on. So what will be the kind of, yeah, test to define, maybe to to solve, maybe to to have some kind of a reference network with some reference problems to be addressed by the by the agents. And this will be particular for the specific mission. If we are thinking on service assurance or we are thinking on checking on the configuration or I mean, there will be probably different roles for different I agents. Yeah. And then we need to define some kind of reference scenario so that we can ensure that the ones that that the different agents basically find the same setup for for performing the action, and then we can have a valid comparison between the execution of these agents in these specific limits of the defined by the by the reference case.

[01:58:43] Unidentified Participant: Okay. Okay.

[01:58:43] Luis: But, yeah, it will be probably not not a simple mission to to address.

[01:58:47] Unidentified Participant: Okay. Okay. Thank you. Thank you.

[01:58:49] Marcus Ihlar: Yeah. Thank you very much, everyone. As you see, there are lots of interesting topics coming in here, and I would urge anyone who thinks that any one of these topics is interesting to engage in the mailing list. That is also to everybody who's proposing things here. Always good to take it to the mailing lists and try to drive discussion there. That also helps us as chairs to gauge the interest in this because there is a lot of work. Okay. So thank you everyone. Hope to see you in San Francisco in a few months, maybe under a new name.

[01:59:25] Giuseppe Fioccola: Just one last remark. Also, Thomas mentioned in the notes, please check that your work is included in the current charter because and also your future plan. So please review the charter. Yeah.

[01:59:53] Thomas Graf: Exactly.

[02:00:09] Marcus Ihlar: I mean, people get awareness. Yeah. They can start the discussions. Right? So Exactly.

[02:00:14] Warren Kumari: Yeah. We release this statement, see that if you can just. Exactly. It's

[02:00:19] Thomas Graf: got already feedback. So

[02:00:20] Marcus Ihlar: Yeah. So yeah. Okay.