**Session Date/Time:** 22 Jul 2026 07:00 [00:00:30] **Simone Ferlin**: Alright. Good morning. This is our ICCRG meeting. Let's wait another two minutes. For me, it's 09:00. The clock there is eighty fifty five. So let's wait a bit more. Alright. More people are coming to the room. We can start. Okay. [00:01:51] **Michael Welzl**: So [00:01:54] **Simone Ferlin**: welcome to the ICCRG meeting on the ITF one two six. I'm going to repeat some information. You probably already know this session is being recorded. This is an IRTF group meeting. So the IRTF follows the ITF intellectual property rights disclosure rules. Some information about that available in the RFCs on the slide. We also record audio and video from the sessions. You have any, objections or restrictions to that, or if you speak on the microphone or want to ask questions during the session, it means also that you consent to appear in the recorded material. Privacy and codes of conduct. Yeah. Routine. Probably know about it. And most important, the goals of the RTF is that we conduct research. This is a place where we present work that is not ready yet or could be eventually forwarded to the ITF for standardization. Therefore, we don't develop standards here. Some more information for maybe new participants, how to choose the your tool, the full or the MeetEcoLite if you're participating remote. Some other resources and agenda, MeetEco information, and if you need some more assistance with technical issues. We are three in this group, so we rotate a bit. Reese is online giving me support from the West Coast. Vidy is also somewhere. And sometimes I'm on-site, sometimes I'm remote. So some people know me here well. The page and the mailing lists or the email from the ICCRG, you can reach us anytime you want. And as usual, we need a notetaker who is going to do us this favor. Someone in the room could do note taking for us, please. Volunteers. No. Yep. Thank you very much. So we have a note taker already. Onto the agenda. So we have quite a few items today, but also enough time to discuss if you want. First, we do the draft updates, and then we go to external speakers and other topics that we bring to the ICCRG because we think they are interesting and relevant for the community. So, yeah, the next person to present oh, yeah. Sorry. So we have also, an update on the active drafts here. So the LEDBAT plus plus document entered the RFC publication queue. Thanks, Gory, for your review and emails that we have been postponing, but things are progressing here. And we have, as you're our first speaker as well, an update on the, pacing. So, yeah, let's let's move on to Michael's slide. [00:06:18] **Michael Welzl**: Good morning. If you know me, you know I can talk long and waste everybody's time, but not today. Two slides giving an update. There are two versions of the draft since the last time this was presented. One is that now the draft discusses ABC, which I think is not really used by that name anymore, appropriate byte counting. This is there's an older version of the document that says also five six eight one, the TCP congestion control RFC. It says that in slow start, you shouldn't increase by more than one SMSS for every act that you receive so that if you lose a number of acts in a row, and then one will tell you, you are now allowed to increase by 10. You know, you that would lead to a burst of traffic, and that is why that rule is in place. However, that makes less sense when you're pacing. And, actually, High Start plus plus specifies that this really shouldn't be done, and we document this. We say that high star plus plus star recommends that, and probably it doesn't make that much sense to have that strict rule in place when you're when you're pacing. Then there is some more FreeBSD text. There was an example that Michael, welcome, put into the document. And that was nice. It had a backpacker drill example, but then he made it even nicer. That's what this is. Just the example is clearer. It highlights the case a bit better. And then we made another update where Vidi finally got the okay internally to give us a bit more details about how Apple is implementing pacing. So the description now explains that there's an API that allows to enable pacing and set the maximum rate. So there's an application cap that is being applied. And then there's a burst the way it's done is that there's a burst budget of approximately two forty four microseconds. It's interesting that this is an approximate number of 200. With a minimum of one MSS, this is kind of I mean, it seems to me to be not very far from the Linux behavior where you have basically one millisecond. This is a different number, but you're collecting for a certain time. And when you have enough of that and then, you know, you exceed that time, then you send that out. So this is a creation of microbursts. Packers are sent time stamp timestamps using a leaky buckets came in that in that case. And these these go further down in the system, just also as in the Linux way that is that we described. And then, yeah, it goes to the AQM. So this is where it's being applied. That is it. We did these few updates because these were things that we wanted to do. There was a request about the FreeBSD one, but no other request from anyone about anything, I think. So, you know, time to have more feedback and more inputs if anybody has anything or otherwise, we could maybe finish that document. That's it for me. Comments, questions? [00:09:36] **Michael Tuexen**: Michael Tuexen, this two forty four microseconds number, I think Vidi said it comes from Linux. She said it comes from Prague. Oh, Prague. So I would be really interested where this number comes from and if someone can give a statement [00:09:58] **Ayush Mishra**: where it [00:09:58] **Michael Welzl**: comes from. That was heralds. It was like Prague. Like, we Maybe you were just [00:10:05] **Stuart Cheshire**: I don't know. [00:10:06] **Michael Tuexen**: Random number. [00:10:08] **Michael Welzl**: It's a Prague number. [00:10:16] **Participant**: Yeah. I mean, there there is an issue there. There's an issue with Prague not so much with pacing as with burst because frag sets the threshold for something ECN marking very low. And so it's fairly easy to exceed that threshold by sending a burst that's too large. And so so that's I I I have no idea what what Apple did, but there is this actual issue in Prague and and burst size. I mean, if if you set your your minimum burst to I think it's two milliseconds in Prague. You you want you want your if you set your queue, the if you start marking the queue with a few milliseconds, obviously, you don't want to send me a second worth of of burst because that's going to interfere. That makes sense. Yeah. Now there is a of course, another issue is that if you want to have high performance with UDP based application that quick, you want to use a general GSO. Yes. And for GSO to be efficient, you want to burst at least 10 packets. Yes. So that's there's a tension there. [00:11:39] **Michael Welzl**: Yes. All true. [00:11:43] **Michael Tuexen**: Yeah. So Michael Tuexen channeling Richard. He wants he he doesn't want to go to the mic. The number is one over 4096, he figured out. [00:11:57] **Michael Welzl**: Oh. So And it's kind of a quarter of a millisecond, but just kind of in line with the two, know, power of tool. Oh, great. Yeah. Cool. Maybe easier to implement with shifting or something. [00:12:18] **Simone Ferlin**: Yeah. So we probably do a working group last call in the mailing list and see what what happens from there given the silence. Yeah. Alright. Thanks, Mike. [00:12:33] **Jana Iyengar**: Thank you. [00:12:36] **Simone Ferlin**: Let's move on to the next update. Windowless. Cumulative. Suck. There, Danielle? [00:12:53] **Danielle**: Yes. Yes? Hello, Chad. Hello, everyone. [00:12:57] **Simone Ferlin**: Good to see you. [00:12:58] **Danielle**: From China Mobile, and my topic is the cumulative SACK extension for RDMA trans retransmission. And next slide, please. [00:13:14] **Simone Ferlin**: Oh, oh, go. K. We want no longer controls. Slide control for me. Okay. Thank you. Not there yet. Oh, yeah. [00:13:42] **Danielle**: Maybe I can share my screen. [00:13:47] **Simone Ferlin**: Has the the controls. So Okay. Let him share. Yes. Now it's here. Yes. So [00:14:10] **Danielle**: Okay. The background is that large scale cross one data transfer scenarios have grown rapidly alongside the rollouts of cross regional computing infrastructure fueled by industrial digital digitization and the cloud adoption. Cross reaching cloud excise and migration carry every large data volumes, creating strong demand for remote data migration. This draft purposes are windowless SACK method aiming to achieve efficient wide area transmission of large scale data over wide area networks through RDMA is the problem is that how to increase throughput under the limited buying wise conditions. So next slide, please. Also, with the development of new Internet sceneries and technologies such as the high dimension video cloud computing, big data artificial intelligence, and the large models, users often need to transmit large amounts of data over wide area networks. At the same time, the wide area network environment is full of uncertainties. So the traditional RDMA methods of go back unpacking loose retransmission is sensitive to package loose, and other retransmission methods rely RTD and window size directly using the traditional packet loss retransmission methods cannot improve the effective throughput. So we we want to enhance the good push is necessary to improve the current SACK mode. So next slide, please. Oops. Next slide. Oh, yes. Thank you. So we propose the dimensional packet loss retransmission mechanisms for timing and accounting based on the network environment and the size of the few packets to be sent. The sending rate is precise, and the cumulative confirmation time T and the cumulative received data packet number n are also determined. So then the window image is removed. The sender doesn't need to adjust the the sending rate according to the packet loss situation. It only needs to continuously send data package message at the precise rate, and the receiver continuously receives data package message within the range of cumulative confirmation. Time t and cumulative received data packet number n, and arrange them in the receiving buffer according to the package sequence number. And after reaching the precise time t or the cumulative received data packet number n, the receiver triggers the confirmation mechanisms and the face back to the received data packet sequence numbers through the SACK message to the sender. So subsequently, the sender checks based on the received SACK message whether there are lost data packets. If there are lost packets, it sends the lost packets data packets to the receiver. So next slide, please. So in order to achieve the SDK mode, we make a new SACK message format and is add a TC_ACKH to the standard RDMA message format. And it's the example of the TC_ACKH message format. Each contents is segment_num. Last finished PSN and segment_1_left and 1_right, segment_2_left and 2_right. So next slide, please. So in this page, we introduce the meaning of the sec of each part and the the rules for the receiver to determine whether to resend the data package are as follows. It's in segment to left to right and so on. Retranspaint the data package number within the range of the segment_1_left last the last the finished PSN and the to the segment_1_left and the segment 1_right to the segment_2_left and segment 1_right and so If segment_num equals to one and the last_finished_PSN equals to the segment one rights and equals to the maximum number of the sent data packets, it means that there are no retransmission is required. And if the segment_num equals to one and the last_finished_PSN equals to the segment one. Right? But there is this is less than the maximum number of the sent data package. It indicates that package lose has occurred. Therefore, data package number from the sec segment one plus one to the maximum number needs to be retransmitted. So this is the the whole process of our draft. Thanks. And any questions? [00:20:35] **Simone Ferlin**: No questions in the room, in the chat? No? No. Okay. Thank you. No. So let's move on to the next presentation, testing congestion control. Mohit, controllers. Thank you. [00:21:16] **Mohit P. Tahiliani**: Good morning. I'm Mohit from an ITK Suratkal in India. This is to give a quick update on the hackathon. I did make the same presentation in CCWG yesterday. So if you are there, you might want to consider finishing some emails because I'm going to repeat the same thing. But, of course, with the focus of ICCRG, as you see this hackathon, typically, we try to do it with the topics that are aligned with CCWG, ICCRG, also TCPM, and TSVWG. We had initially five objectives, but we reduced it to three considering the time that we had in the hackathon. This time, it was relatively shorter hackathon than every year, every every idea that we have. So we had three objectives. The first two are actually aligned to TSVWG. They are on queuing. The third one is related to the PicoQuic, which is a quick, you know, integration that we are done with n s three. And then I have some more points to discuss that are pretty relevant to ICCRG and the and the Internet drafts and the RFCs that have been produced at ICCRG. The first objective talks about the comparison between FQ codel, is RFC eight two nine zero, and FQ-PIE, which is, our Internet draft in DSPWG at this point. And, we did the comparison of both of these algorithms, using a Samsung phone, configuring these, algorithms as the mobile hotspot in the mobile hotspot, and also in a WiFi AP provided to us by Quantum Networks. So thanks to Samsung and Quantum Networks for giving us the devices. This was the setup. So we had a client. This is my student's client, Abhijar, who is on-site with me, and then there's another student, Vishal, who participated remotely from India. So this is his laptop, which was connected to the mobile device, connected to the access point, and then we were using the buffer bloat server. I'll not spend more time on this. As I mentioned, this is more relevant to TSVWG. But if you are interested, the slides are up. The Internet draft is in TSVWG. If you have some points to discuss, I'll be here throughout the ITF. We'll be happy to discuss these points. This was the setup that we had for the access point. Pretty much similar, just that instead of connecting to the mobile hotspot, we were connecting to the WiFi AP. Again, we had results done for this. For running these experiments, we used flexible network tester, which is kind of a Python wrapper on top of iPerf and other networking tools. It has implemented some specifications like r rule, which is real time response under load. So it pumps in bidirectional traffic and then uses ICMP packets to see whether RTP is acceptable or not. It also has several other tests. We chose some of the tests to run during this experiments. I'll go to the next slide with just a observation that has stood out from last four ITF hackathons that we have done is something that tail latency in FQ-PIE we see to be better than FQ-CoDel. We are still running more experiments. We have done more experiments on NS three, but these were the experiments on Wi Fi. The ones that we did on n s three were not using Wi Fi. So in all different setups that we have seen, we see a clear improvement in the tail latency. But, yes, the work is still going on and more experiments are to be performed. Talking about now something related to congestion control, which is also relevant to CCWG and ICCRG is PicoQuic. So in past few ITFs, we noticed that PicoQuic is something that a lot of people were using at ITF for different kind of work. And also, we saw recently according to the stats from last five years that more than 1,500 papers have been published using NS three. But NS three does not have a native model of QUIC. On the contrary, we have several other user space implementations of QUIC that are available. So we decided not to build a native model in n s three, but rather to use something that ITF is already using and integrated directly with n s three. So PicoQuic is a c implementation of quick and multipath QUIC. Christian manages it very actively, follows RFC 9,000, 9,001, 9,002. And on the other side, NS three is something that provides you reproducibility, scalability. So we are now integrated it. This is a sample result. If you are interested to look at the code, here is the QR. You can scan it and you can look at the code. It's a very thin wrapper that we have built to integrate both the libraries in such a way that if one goes on to get updated, it will not impact the integration as such. So you can actively keep working on PicoQuic and keep adding your algorithms or your features. You can also continue to work on n s three in parallel and you wouldn't need to touch this wrapper again unless there are significant changes in the public APIs that PicoQuic exposes. We have also provided around five to six examples for the users to get started with. One example is competing congestion control. So Christian has an Internet draft in CCWG which is called a BBR. You can see one experiment that we have run, has a flow with BBR. There is another experiment that runs cubic. You can use BBR. You can use other congestion control that PicoQuic supports. And you can use these several topology helpers that are available in n s three to quickly spin off a topology and then run these experiments using PicoQuic. So these are, the works done and we have also provided one example for multipath QUIC. NS three did not have any example of multipath even not even multipath TCP. So there's a first multipath example that now NS three will support. So please feel free to use it and do let us know whether you find any issues or if there are certain things that we missed out. Something that we did not do at the hackathon, but we did between the last ITF and this ITF. The first bullet is related to the work happening in TCPM on PRR. So now we know PRR is a standard, in 09/09/1937. The earlier one was an experimental, RFC six nine three seven. And, we found a small bit of alignment that was supposed to be done in FreeBSD. So a couple of my students worked on this project, and we pushed their patch to FreeBSD. And on Saturday, I met Richard and Michael, and thanks to them, now that patch is merged. So this was the first contribution from our institution institution to FreeBSD, and we look forward to having more active participation. The next bullet talks about two things that are directly relevant to ICCRG. LEDBAT plus plus, I saw Simon's slides. I think now it's in the editor queue. But our LEDBAT is already RFC nine eight four zero. There was no implementation of these two algorithms in NS three. So our team of students have implemented LEDBAT plus plus and rLEDBAT, these are the links to the merge request on the GitLab repository of n s three. LEDBAT plus plus has been already reviewed by the n s three community. I'm a maintainer for TCP and queue management in n s three, But then there are other n s three maintainers who are actively looked at LEDBAT plus plus and given some very good feedback, and it's on the verge of getting merged now. But we haven't yet received much feedback on our LEDBAT, so I would encourage community members if you are working on any of these algorithms and you plan to test them or use them, it would be great to have your feedback on this implementation. For example, prior to implementing LETBAT plus plus and rLEDBAT, LEDBAT was already available in n s three. But when we worked on LEDBAT plus plus implementation to get started, we identified there was a bug in the LEDBAT implementation of n s three. So that was a pretty serious bug and that went on to get fixed now. So we had to fix that and then implement these algorithms. So more people looking at it will definitely help us to ensure that we have lesser bugs in these implementations. Apart from that, like you saw, I mentioned in my previous slide that we used Flet to run the test. So what we noticed is lot of congestion control algorithms and queue management algorithms are supported in NS three, but we don't have an application like Flet that people could directly use. So we have built now as a part of the Google Summer of Code project, the student who is working with me in this year. We have built a Flet like application for NS three. So if you had a r rule test that you are running on a Flet on a live test bed, you could use the same r rule test now in an s three and it'll give you a similar type of results in the same format, same candle plots, same cumulative distribution function plots. So we'll try we'll try doing this implementation in a way that you don't see the difference whether you are running on n s three or you are running on a live, test bed. The code is here. This project is still going on. The student is going to submit the final code by the end of this month, and it will be then available on the GSOC website for review. So please take a look at it and let us know if you have any other observations that we should improve on. And finally, we have added some new topology helpers in our tool that we have built at NITK. We call it as network stack tester. It is a Python wrapper on top of the Linux network namespaces that allows you to quickly spin off topologies. The most important topology helpers that we have added are for multiple bottleneck topologies. So dumbbell topology is very well known and used for congestion control, but we noticed that now there is a lot of interest in multiple bottlenecks also in the network. So we identified a couple of papers from nineteen eighties called generic fairness configuration one and generic fairness configuration two. These are two very well known topologies for evaluating congestion control on multiple bottlenecks. Now we have these two helpers readily available in Nest. You can quickly spin off a multiple bottleneck topology without having to spend time. And those are already merged in NEST. Here is a code and some examples for that. So that's all from my side. What I wanted to mention is that we will continue to build more features and more algorithms in NEST three or NEST. What we seek inputs from ICCRG is to let us know whether, if there is any protocol of interest to ICCRG or if there is any feature that somebody would like to see so that we can allow more Internet drafts to be evaluated in the process as they become RFC. Thank you. That's it from my side. [00:31:58] **Simone Ferlin**: Questions from the room? We have been in touch with you, especially, Reese, about having the like, a a similar yeah. A a similar evaluation suite that you have for AQM but for congestion control. [00:32:20] **Mohit P. Tahiliani**: Is this Oh, yes. We have so I did not mention about the so we are we are building a congestion control evaluation suite, which is in compliance with RFC nine seven four three. Mhmm. We have made very good progress on that, but that still is work in progress. So maybe in the next [00:32:39] **Simone Ferlin**: It would be nice to hear about it. Sure. Yeah. [00:32:42] **Mohit P. Tahiliani**: Okay. Thank you. Alright. [00:32:45] **Simone Ferlin**: Thanks, Mohit. So our next speaker. [00:33:00] **Ndumiso Gasela**: Thanks, chair, and thanks, everyone. Good morning, and apologies for not for not being with you on-site. The presentation is based on Little's Law for queuing theoretic approach for congestion window update. The easy part here is that this is all about congestion window update. We have addition additive increase, multiplicative decrease, and here and cubic increase. Here, we are looking at your at Queueing theory to determine the congestion window update. Would the the the presentation is based on a paper that has been published in electronics, m d p I, and another paper that will be published soon, and I can share that once it's published with everyone in the main list. Let me see. Okay. The yeah. So this is sort of the problem we're trying to address, which is a common problem that we we always address when we do congestion control algorithms. It's about throughput delay efficiency, especially to ensure that we can maximize the power metric. We have, on one hand, high utilization, high delay regimes. On the other hand, we have got the low delay underutilization regimes. We want something that's balanced so so so so that we can improve the power metric. Okay. Just a comparison of the different approaches. The first one is around loss based, which exhibit high standing queues and loss cycles. We have got delay based, which tend to be to under utilize the the bandwidth. And then we have got b b r, which is a model based, and it tries to implement clay Clair Rock optimality to ensure that the in flight data is equivalent to the base bandwidth. What we are trying to do is also to achieve the same, but by regulating RTT using queuing theory. That requires a suitable RTT target, which probably is the biggest problem, but we that's something that can be solved separately. Queuing is based on the way the sender on the end or the or the end user looks at the network. The sender looks at the network as a big pipe or a queue. So we're looking at the network as a queue and we model it using Little's Law. Yes. Indeed, there could be argument to say, you know, the the network has got a the topology can be very complex. But from the sender's point of view, they simply the sender simply sees one pipe, and that's what we are using. That pipe can be stochastic. And in this presentation, obviously, we do assume static, not because we can't also do computation when the pipe is assumed to be stochastic. In in that way, then we use Little's Law, which says that the occupancy is equals to arrival rate times delay or propagation delay. Typically, Little's Law work on averages, but, of course, it assumes a certain amount of time, running time to average, and that running time could be milliseconds, could be microseconds. As long as you can do those measurements and it's reasonable in the environment, you then can be able to apply Little's Law. But also if it's possible differentiate the occupancy, to differentiate the delay, which is r here, then we can create a differential relationship of Little's Law. And in this case, that's what we are doing. And we are assuming that if the delay does not change and the delay the delay remains constant, then we can have a condition. We refer to this condition as stationary delay. And that condition says that the change in occupants with respect to arrival rate should be equal to the delay. Obviously, as we move towards high utilization, the delay becomes so high, and we we attempt to keep the delay at some at some level that is specified. And in that way, we define an r target, which is defined by alpha times the propagation delay or, in this case, proper propagation round trip time. So under this adopted single bottleneck, the condition then motivates us to operate near the base bandwidth delay product, which is clean rock optimality. And it also allows favorable throughput delay trade off. Based on the condition, the stationary delay condition, we are able to develop or design a predictor step and a corrector step. A predictor step is trying to estimate the deviation from equilibrium or the target RTT, the target RTT. So the predictor step is trying to to calculate that deviation, and the corrector step step is trying to find the congestion window that will counteract that deviation. So we first measure congestion window w, in flight data, and occupancy, and and as well as delay or round trip time delay, which is r. And then we predict the the the the influx data that we'd like to have in order to ensure that there's no deal deviation and we correct using a congestion window, and we use that in a cycle. Obviously, to do this, we also while while the the purpose is really to see how we can use these relationships to do congestion control, there are practical issues in the implementation. The first practical issue is around minimum RTT tracking. Currently, we are using periodic probing probing that is similar to BBR. So we are able to update minimum RTT whenever RTT is lower, but also to detect minimum RTT when it's increasing. Also, you have to make sure that there's fairness and coexistence because any congestion control algorithm that doesn't have fairness, then it's no point even to start. There's no viability with that. So, again, the RTT probing does does help in terms of fairness, but it also reduces bias for for stale minimum RTT. And in that way, it it it it improves fairness. We also use low stats like in in New Reno or other cubic to start renew new stat to I mean, sorry, slow stat before we get to con to congestion avoidance. In this case, when we add con congestion convergence, we then use QTSDR, which is the just we are denoting the the algorithm that we are developing, which stands for queuing theoretic stationary delay regulator. The default alpha for our RTT target is 1.2, which we have determined empirically to be the well, more or less stable, but also low enough, which is about 20% of minimum RTT. Then in terms of the simulation, the first one we have basically done a single broccoli neck as simply topology, and we are assuming 100 megabits per second, one way delay of twenty mill milliseconds, which translates to base round trip time of forty milliseconds. We ensure that the paper the buffer size is large just so that we we we can see that we can regulate the the the the target RTT. So the first is about regulating or optimizing the in flight data so that it's close to the base base bandwidth delay product. And in this case, you can see QTSDR is very close to it. BBR normally is a little bit high. Higher, it's always about two times the the the the the base BTP. Obviously, that that's a design issue. Then loss based algorithms are quite high in terms of in flight data, and they depend on the buffer or the the the buffer size. This is the distribution as far as the RTT. As you can see, we are first able to reduce the RTT so that it's closed to forty milliseconds, which is the base RTT. And, also, the the the oxygenation is low, which we can see around the interquartile range shown by the post plot by the post plot. The the next one is the power metric, and this is normalized power metric. And what you see is that the QTSDR is quite high, close to 0.8 in terms of the parametric compared to cubic and neuronal, which are quite low, around point two, and BBR around point five. Then in terms of the buffer size, the effect of the buffer size, there's no effect in QTSDR. There's no effect on BBR. And then on cubic, higher the buffer size, the lower the power metric. So, again, we are able to ensure that the algorithm operates independent of the buffer size. Then in terms of fairness, maybe before I touch on fairness, the slide that I didn't put here talks to good put or throughput. What we find that QTSDR is comparable to all the other algorithms. So the fact that it manages or it regulates RTT, it does not then make it deteriorate the throughput. And then in terms of fairness, again, the fairness is comparable. All of it is above point nine, which is reasonable fairness. Again, we are saying is that the algorithm does not deteriorate fairness. Of course, this is a fairness of the same type of algorithm. We are we haven't done a furnace in heterogeneous configuration. So in conclusion, one QM theory provides a practical RTT regulation principle, and we have shown that stationary condition can be translated into a set aside predictor corrector condition window update. We have also evaluated the QTSDR, and we show that it remains close to base. BTP BTP substantially reduces RTT and achieves higher normalized power. And then the results, it shows that when we can almost complement loss response and bandwidth prob probing. Then the next step is looking at adaptive target selection of the the the the the alpha or or the target RTT so so that you can dynamic dynamic regulate it dynamically. And then last point that I want to add then is that, you know, we don't see this to necessarily becoming a stand alone and probably mainstream congestion control, but we see the concepts that can be used with other congestion control algorithms or even motivate future approaches in terms of congestion control. Thank you so much. [00:47:16] **Simone Ferlin**: Go ahead. Has it turned off? Can you use this? No one in the front. [00:47:36] **Participant**: Thanks for the presentation, Roland. One thing is that, typically, it's not that trade off that you showed in the beginning. So you can't have low delay and high throughput, actually. So that's just one comment. So it's it's not mutually exclusive. The other thing is I'm I'm not sure how I mean, curing theory is nice, but, actually, the systems that we are using are in this row equals one regime, which means, typically, then you cannot, like, apply queuing theory quite well. But it seems to work somehow. So so my concern is, did you test this with, let's say, loss based congestion control at the same time? So typically, typically what we see is that delay based congestion control gets suppressed by loss based congestion control flows. And the other thing is low delay based congestion control typically has problems with varying, let's say, delay due to access media, for example. So did you test something on Wi Fi? [00:48:50] **Ndumiso Gasela**: Yes. No. Thanks. Let let me just quickly start with the the loss. Yeah. The coexistence of loss based with that has not been tested. However, the let me just give the belief. The belief is that what will happen is that it it will then also start increasing the the the the the base RCT with based on the increase with the loss based. Obviously, that could potentially cause a bit of instability, but but I think that's just a shortcoming of of the approach. Yeah. So we haven't tested it, and the belief is that because of the ability to detect the increase in RTT, it will increase with loss based RTT. Of of course, that is not desirable. Then on the delay based congestion control, the this if if I suppose it has the same sort of problem that you'll typically have in in most delay based congestion control. However, we have tried to mitigate that by ensuring that we can quickly detect when RTT increases and as I explained with the loss based algorithm. The the main purpose right now is really to show the mathematics that it can work. Again, the mathematics, as you rightly say so, it depends on Licklaus law applying and, yeah, the conditions to ensure that Licklaus law can apply. That is the the the the the environment is close to endogenous as opposed to endogenous. [00:50:50] **Participant**: Okay. So my main point was you you should test it with with, let's say, artificial jitter or natural jitter that is caused in the system. So so that could the variability in the in the delay measurements could actually be challenging for your algorithm. [00:51:09] **Ndumiso Gasela**: Yeah. No. Thanks. Yeah. That that's true. We we I I agree with you there. Maybe I can just add to say that we we I will I will put this on the git on GitHub. And probably if that that is the implementation on NS three. I'll put it on GitHub for anyone who would like to test it. And yeah. Yeah. I don't know if it's possible, Simeon, to put it also on data tracker. [00:51:45] **Simone Ferlin**: On the data tracker? No. [00:51:51] **Ndumiso Gasela**: Okay. We [00:51:57] **Simone Ferlin**: don't have more questions in the room, so let's move on to the next speaker. Thanks, Domisa. [00:52:05] **Ndumiso Gasela**: Thank you. [00:52:11] **Simone Ferlin**: So hello, Minqing. We can see the slides. We can see you. [00:52:18] **Minqing Wu**: Yes. I can you hear me? [00:52:21] **Simone Ferlin**: Yeah. Thanks for accepting our invitation to come by and talk about Mooncake, you can take it away. [00:52:32] **Minqing Wu**: Yeah. Thank you. It's my honor to share our progress on disaggregated large language model inference in this ETF conference or or workshop. Today, I will mainly talk about our rest and the progress on how we improve the transcription that would which is a core component and of the home care existing. Also, we are talking about some background such as the new agentic workload and why it is important. So, you know, that currently, the large language models are I'm sorry. Some some pictures are missing here. But, essentially, currently, the applications of large language models have changed from the original chatbot, which typically we have only a single term response to a genetic workload, which means that you will use a harness that will driving a margin term and very complex execution topologies, which also means that it will take a lot of tokens. It will consume a lot of tokens per per request, which also leads to the increasing of the consuming of the inference cost, which, as a result, the current GPU are mostly used in the in inferencing the models rather than the training the models. But, originally, we only built in the so called compute infrastructures, which is typically measured by how many computation power such as the p probes, how many petabytes peta probes that this computing infrastructures can hold. But nowadays, we are we want to build not only the compute infrastructures. We want to build the token factories, and a token factories should be measured by how many tokens that it can consume, how many tokens it can produce per second. And there is a very huge gap between the computation power measured by data flows and the the token processing capability measured by token pass per second. And the the difference or the gap is determined by the your inference infrastructures, And that is the reason why we need to build such as the the title of this slide, the disaggregated inference architecture to improve the the capability of transforming the computation power to the token processing power, and that that is the main topic of of today. And, of course, that as you can see that in our traditional computer computer architecture, we will build some kind of hierarchical of memory that which means that we will have also a very expensive, very small capacity, but very fast, the the the catch, and we were combining it with the slow but more cheap SSD over DRAM. We will combine a hierarchical of DRAM in order to balance the the cost and the bandwidth that we can achieve. Similarly, in the intelligence or in in the nowadays, large language model process, we can also have different kind of heterogeneous accelerators, such as we have high end GPUs, which equipped equipped with high bandwidth memory, which is expensive and small. And we also have middle layer GPU accelerators, which is equipped with GDDR, which is also which have a slower bandwidth, but maybe less expensive. And we also have DRAM, which is the cheapest in terms of the capacity per dollar. So if we can combine all these different kind of heterogeneous accelerators, we may be able to have a better TCO with with of our super accelerators. And this is the reason why we will this is a more detailed explanation. For example, in I'm from China, and then in China, we have several different kind of accelerators such as the h 800, which is an flagship accelerator. So it is typically good at t forks per third. So it is computation power is good. And, also, we have h 20, is the limited, which computation power is cut it because of the limitation of exports, whatever. But it is also more cheaper, so it is good at the bandwidth per dollar. And, also, of course, we have the CPUs equipped with DDR five, which will have which will have bad at the which will be good at the capacity per dollar. So if we can desegregate the AIM inference into three stage, One stage consuming computation, one stage consuming bandwidth, and the the third stage consuming capacity. We can use different kind of accelerators for different stage, and that is the main idea of our, Mooncake system. So, essentially, this Mooncake system is, they are it is based on the, internal character characteristic of the large language model inferences. As you can see that, it can be partitioned into two stage. The first stage is called the pre fail, which will consume all the input tokens simultaneously. It will process it then concurrently, so it consume the computation power at most. And then in the second, the decoder phase, it will decode one token per iteration. So it for every iteration, it will scan order it reads of your large language models, which is very huge. So in the the code phase, it will consume the bandwidth at most. So now we have two stage, one consuming computation and the the other consuming the bandwidth. And how we connect these two stage? Actually, we connect them by transferring one key trans one key metrics called k v catch from the previous stage to the decoder stage. And this k v catch, it is called catch because it can be catched and then reused in the second round. So as you can see, if you have two different questions, one question is, what day is it today? And then the second question is, what day is it tomorrow? As you can see, the first four tokens are the same in these two questions. So they can we can preserve the original Kiwi catch, process the output result of the first question, and then we use the KiwiCatch for the first four token in the second round. So this new order with this reusing, we can reduce the computation by storing and then reusing the KiwiCatch. So this storing and the reusing, which means that the catching of the Kiwi catch will create capacity at the most. So now we have three stage, one computation, one by device, and the the third is full capacity. And we can as as we have mentioned before, we can use h 800 good for computation for preview. Use h 20 good for bandwidth for the code, and use d DDR five good for capacity for catching all the reusable KiwiCatch. And that is the main idea of our disaggregated architecture of large language model inference, which is called Mooncake. And the list Mooncake system is originally the also the currently, the inference platform of the Kimi, which is known or in the last year, it is known for it is open source model k two. And in the just two weeks before, it is also used for serving the current most open week large language models, Kimi k3. And we will open source Kimi k3 in open week Kimi k3 in the next week. So it is it has experienced very high bandwidth and the high very high concurrent user usage. And it is main benefit is that we increase the throughput of Kimi by as large as 75%, which is by using this disaggregated architectures. As you can see in the right side of this slide, we have a stand alone prefilled instances, which is consumed by a a computation, no intensive GPUs such as h 800. And we also have a standalone decode pool of GPUs, which is consisted by bandwidth intensive accelerators such as h 20. And then in the middle, we have the KiwiCatch port, which is appalling all the DRAM and the RDMA bandwires of all these GPU servers. So we have three different con component. And more importantly, because currently, the large language models are all sparse and OE models. So in different for in different pre field and in the the code, typically, we will use different different parallelism strategy. For example, in the pre field, we may use tensor parallelism. And in decode, we use expert parallelism. So these different parallels parallelization also require that you need to desegregate the preview post and the the the code post. And in order to connect all these different component, you need to take make best use of your GPUs, and that is what we have open sourced the Mooncake the the Mooncake lab open source library, which is useful. And that the second, the goals of this Mooncake open source project is that we we have we we have contain all the reusable KiwiCatch in this very large capacity KiwiCatch port. Because currently because if you have a lot of concurrent users, you may the KiwiCatch per user may be several megabytes, and when it multiplied by the maximum number of users that you want to you want to sell, the result aggregated size of the Kiwi Catch size can be hundreds of terabytes or even petabytes. So this is a very huge size of the KiwiCatch. And with all these at least these are the original usage of the mooncake, which is used for catching catching the Kiwi catch and also transferring the from the preview to the decode side. But from from three years ago, the Mooncake system has developed in many, many more usages. Currently, we are not only disaggregating the large language model inference to only three stage. We are actually as you can see in this slide, we have disaggregated into five or even six or even seven multiple different components such as, for example, currently, the large language models not only not only contrast the text materials, it can also contrast images and the videos, which is processed by a standalone component called a vision encoder. So we can disaggregate it at this vision encoders to some more GPUs, and the list list to another desegregation architecture called EPD, which originally is prefilled and the decode PD desegregation. Now we have we should prefilled the code desegregation, so three component. And, also, originally, we only use desegregation in a single data center. But now we we have proposed an architecture called prefilled as a service, which means that you can keep the prefilled GPU on one data center and that the decode GPU on another data center and connect them via some different kind of via the the wide area network, and that is possible. So we have evolved the Mooncake system into different component. Once you want the disaggregation, you will you you can use Mooncake to connect the different components, and that is the main of goals of why we developed the Mooncake system. And, also, the Mooncake system has been open sourced and used in many companies such as the not only in the Moonshot AI, but also Alibaba, Vocano engine in approaching AI and finance and also Nvidia and many many many component Internet components. And and it's also integrated in most in in or most of the open source, numerous engines such as vLLM, SGLang, and NVIDIA's Dynamo system. So it is one of the most standard, the middle view for the current large language model inferences. And the the main objections, as you can see, we want to transfer the KiwiCatch or other materials as fast as possible, and we want to store more KiwiCatch in order to reuse them as much as possible. And we also want to keep it easy to use so that you can use it in the more and the more scenarios. And that is the main objections of Mooncake. We can because of the time, we we I want to directly go to some our recent advance of the current Mooncake. So as you can see that this is the original experiment result of how we use Mooncake. The x axis is time between tokens, which means that how many tokens can be produced per request. And the the y axis is the capacity throughput capacity. As you can see, at the lower time between token you want or in in another term, the la the the larger the tokens per request that you want. For example, if you only want 10 tokens per second per request, you can you do not need the disaggregated action. But you if you want 50 tokens per request, you need to disaggregate the preview from the the code in order to have a better throughput, and that is the main reason why we use decode Monkey. But now this is the original case. But now we found that the current network is evolved into and very a more more more and more complex, which means that we can have many heterogeneous network interconnects. For example, we if you direct the RDMA for steering out network. We have only link in the scaling up network. And, also, we do have for the core data center network. But in the original, this is to a so called a multi class selection program. But in the original design of the Mooncake, we need the users to improve imprint you to select which protocol you use before you're using the Mooncake transfer engine, and that is not good for this kind of of heterogeneous network process. As as you can see that originally, we ask users to impartially select me which kind of network protocol you need at the start up time, and the list is statically binding. And then we will have a fixed and a state planned past selection and the scheduling policy, which will lead to a very large P99 latency because of the congestion of your network. And, also, if some of the network interface card that is failed during the the transferring, you need to recover the network transferring manually, and that is not good if your clusters the size of your cluster is very large. So in order to resolve all these challenges, currently, we have we have developed a new technology called transferring engine new technology. And then it is it will want it the the main objections of this new technology is to resolve this all these three problem. The first is when we have originally we select to use GPU direct RDMA, and it is failed. We want to use some kind of relay in order to recover from this kind of stack binding. And we also want to recover from the congestion, which means that we want to rebalance the network in order to porting all the network interface card in order to make best use of the aggregated bandwidth. And, also, we want to automatically recover from a a failure a partial failure of only one network interface card. Because, you know, currently, for every GPU servers, you do we will have multiple multiple network interface card. So if only one network interface card is done, we can still use other network interface card to recovery the to to recovery the connection as long as we can relay it, for example, via the a meeting. So with all these three objections, we have developed the network transfer engine new technology, and it is main objections and abstractions is that we use a unified segment abstractions to unify all different kind of network topologies. As you can see, because that essentially, all the large language model inference workarounds are essentially moving some parts, some byte buffers from some locations to some other locations. So we can abstract away all the different kind of protocols by using this memory layout. We just pro we just provide the size, the pointers, the offset of some byte buffers, which is the source byte buffers, and also give the destination byte buffers. And we just declare clarity a declare declaratively is saying that this is my network transporting intentions, and that it is the responsibility of the monkey transfer engines to choose automatically which network protocol we use, which which route we use, and that that can be all handled automatically. And it can be a meeting. It can be RDMA. It can be a TCP, and then it can be even TCP over wide area network. And that and the the material can be DRAM, can be HBM, and it can also be SSD and HDD because we can also abstract SSD by by the buffers, and they can be all the different kind of implementations can be all abstracted. We we are this u unified abstractions. And then we can choose different but even without the difficult the different implementations, we still need some kind of locality affinity. So we need to automatically finding out what is current topology of the network in order to help the system to choosing the locality affinity. We try best to do not cross data center, do not cross PCIe switch, and do not cross the NUMA selection. So that this is automatically found and used in the dynamic route selections. And then we can also install different kind of protocols by by even with the same obstructions and the then even if, for example, we currently have RDMA available, NVLink available, and the TCP available. So in the first choice, we will choose RDMA. And if the RDMA is not usable, we can use NVLink as a relay to find other RDMA route. And if both the RDMA and and the NVLink is not usable, we can still fall back to the TCP link in order to make our transferring as high available as possible, and that is very critical in the real world log language model transferring. And because of this kind of heterogeneous network, we need some kind of spring splitting of the network package in order to achieve the best latency. And the list is what we call the adaptive slice splitting mechanism. Essentially, it is very similar to the network breeding of the traditional TCP packages. We just have some kind of latency estimation including by measuring the network inter in flight bytes, the size lens, the penalty factors on the the bandwidth, the current usage of the bandwidth. And we use all these we use all these metrics in order to achieve a very fairness the fairness of the network. We choose between the network interface card, which is the most low which network interest interface card can currently have the lowest latency, the lowest band bandwidth usage. So we choose this route in order to have a better aggregated ban wise. And, also, this all this aggregated ban wise and aggregated appalling of the network multiple network interface card, we can also assure the resilient under the sea of hearing in order to achieve the net resilience, which, as you can see, just the the link level resilience. If one network interface card is is failed, we just find the other network interface card that we can still use to connect with each other. And even if the RDMA is fully down, we can fall back to TCP cases. And, finally, we have we can currently present some evaluation of this network interface, monkey transfer engine, and new technology by comparing with our original implementation, the monkey transfer engine, and the immediate implementation, the immediate mix source. And as you can see that, essentially, because we can dynamically choosing the best network interface card, when we are transferring a large chunk of data, they it can have a better throughput or a lower latency. The the reason why the larger chunk we have, the larger throughput again we can have is because that when we transfer in this large chunk originally in the original implementation, it is statically partitioned up to 64 kilobytes and evenly spread to multiple network interface card. And it is typical some of the network cut will be slower than the others. So this is struggle. We are lead to the average group put lower than the expected bandwidth. And by dynamically selecting the network interface card, we can have better usage of the of the bandwidth. And, also, we make use of this better bandwidth in other many real world usage, such as, for example, we use it in the reinforcement learning. In reinforcement learning, one important case is that after the training, you need to update the newest version of your model weights to order serving inference in order to load out. And the list needs to transfer the whole weight of your of your model. And currently, the size of the model is larger and larger. So this this chunk is very large, and the the the transfer engine new key new technology is better in this case. And, also, we we have also used in the KiwiCatch reuse scenario where the larger size or the longer context you have, the the better benefit that that is the transference in MTE will bring you. And, finally, as you can see that all our technologies have been open sourced in this URL. You can found this organization in the GitHub, and this Mooncake is the host rapport host or the source code and the evaluation reports that we have just presented. So if you are interested, you can come to the GitHub, and we we can if you are interested, we can also collaborate it on this open source project. So so that that's my presentations, and I'm happy to take any questions. [01:19:33] **Simone Ferlin**: Thanks, Minqing. Any questions from the room? No? So the one of the motivations that we invited you here is also because we use Mooncake quite extensively at Red Hat to look how LLMs work as workloads, as applications, as we develop the platform and also integrate the software with these different accelerators. We had, at the beginning of this exercise, very little idea how all these operations in LLMs from PD disaggregation, KB cash exchanges, very little idea how these things behave on the wire and whether we are using any libraries provided by NVIDIA that integrate very tightly with their hardware and run over RDMA, or we use, like, a distribute an engine for distributed inference such as LLMD, which is something that Red Hat is behind to be agnostic to the transport, whether it's RDMA or TCP that is underneath. So we have been working with TCP zero copy to see how far we are in terms of performance to RDMA. So zero copy is something that, if I'm not mistaken, Google upstream is not a long time ago. We back ported. We work a little bit on that. So the main motivation here was to have Mooncake as a is basically a a trace basis application that you can use to understand how LLM as an application behaves on the wire and how the network can or cannot absorb all these operations that the these models do. So thanks very much for for coming by and showing your work. Your paper and your open source stuff is very useful for us at Red Hat as well. Thank you. So let's move on to our last talk of the day. Ayush. So [01:22:11] **Ayush Mishra**: Okay. So my goal for today is to try to convince as many of you as possible that maybe bandwidth fairness is not the right goal for the networks that you're running. And I'm going to try to make this argument by presenting this paper that we presented at NSTI earlier this year, where instead we said that what we call approximate performance isolation is a better goal for for for your networks. So if you run a network, the congestion control algorithms that run within your network are going to be very important. And if there's any place where this is not gonna be a hard sell, it's probably ICCRG. And this is not just because of the reasons of performance, but this is also because it's a it's consequential in terms of how you size your buffers, what EQMs you deploy, and in the context of the Internet, especially, you know, how you think about fairness and the deployability of these congestion control algorithms themselves. So this last concern here, which is regarding fairness and deployability, this has gained particular attention in the recent years, especially in the Internet context. And a lot of this is driven by the fact that today, the Internet hosts a heterogeneous mix of congestion control algorithms, often a mix of algorithms that have been found not to play very fairly with each other. So for example, two of the most popular congestion control algorithms on the Internet are Cubic and VBR, and there's been vast amounts of work that has revealed that these two algorithms don't, play fairly with each other when they interact with each other in certain network conditions. So just to quote one of these measurement studies for 2019, they observed that, you know, a single BBR flow was consuming a fixed 35 to 40% share of the bottleneck bandwidth while sharing the same length with as many as 16 cubic flows. Now, of course, given the significant, unfairness results that have been found in control tests, there has been a lot of, discussion on how we can make these congestion control algorithms more fair, how we can review the deployability standards on the Internet, and so on and so forth. These discussions have happened both at Sicom as well as ICCRG and CCWG here at IETF. Now what I find particularly interesting about this development is that we have been doing congestion control on the Internet for decades, yet this heterogeneity driven unfairness on the Internet seems to be pretty decent. And I sort of wanted to delve a little deeper into this question to sort of understand where this unfairness is really stemming from on the Internet. So to answer this question, what we did was about in 2024, we built a measurement tool that can measure what congestion control algorithm a website is using as well as over live sessions, figure out if the video stream is using a different congestion control algorithm compared to your ad loads or something else that the website might be doing. And some of the key results that we found out was that, you know, different websites deploy different congestion control algorithms. But more interestingly, you know, these algorithms can differ across regions. So different regions the same website can deploy different algorithms in different regions. And even within a website within the same region, depending on the content that you're delivering, you could deploy a different congestion control algorithm. So for example, over here, on twitch.com, we saw that the video was being delivered using VBR while the static assets and the banner ads were being delivered using Cubic. So this really makes things very interesting, and this really points out that the heterogeneity on the Internet is really driven by the fact that a lot of people are trying to do a lot of different things on the Internet. More specifically, I think what this development brings into question is this really this notion of bandwidth fairness that we have been pushing on the Internet for many decades now. And the reason I want to review this question is because bandwidth fairness is a great goal for the Internet if two things are true. The first is that your bandwidth must be finite, and this does make sense in the Internet context. And it makes sense to come up with a good way to share bandwidth. But the second stronger assumption that bandwidth fairness comes with is that everyone only cares about throughput. Now given these measurement results, the second assumption is sort of weaker, today on the Internet because if everyone did just care about throughput, they would all be deploying the same congestion control algorithm that optimizes for just throughput. And the reason this doesn't happen sort of falls in line with the observations that I made earlier, which is that because different people deliver different things on the Internet, they want to make different trade offs, and therefore, they choose different congestion control algorithms. Now this is not a new result. This is something that we know really well. Right? So for example, two congestion control algorithms in the Linux kernel. Let's pick Cubic and Vagus, for example. You can see that they use the same network very differently. So for example, over here, Cubic, which is a well known buffer filler, it's gonna fill the buffer packets in order to maximize throughput, whereas biggest is gonna be a lot more conservative, give up a little bit of throughput, but minimize delay. Now the problem in the unfairness happens when we try to mix these competing goals together. So if I take the same two algorithms that are performing reasonably well here in this isolated setting and I put them together, we see that, you know, now Cubic seals most of the bandwidth and Vegas stars for any any bandwidth necessary. In particular, what I would draw your attention to over here is that in this two dimensional figure of throughput versus a delay, what's essentially happening when these competing objectives compete with each other in the network is that an algorithm shows significant translation from where it wants to operate to where it ends up operating, under contention. And this is not something that's unique to just these two algorithms. It it's true for most algorithms in the Linux kernel that have been deployed on the Internet. So we have pulled four more algorithms on the Internet where all of them make different trade offs. But when I put them together, again, you get the significant translation from where someone wants to operate to where they end up operating. Now, again, this is not, breaking news. We all know this really well. So what am I really, giving this presentation for? So the main thing that I want to motivate via these observations is that because these algorithms represent these different trade offs, a good goal for allowing these algorithms to coexist is not just bandwidth fairness, but maybe trying to minimize this translation between where they operate and where they end up operating when they are competing with each other. So, essentially, for the Internet, I think what we should try to, I mean, try to keep as a goal is make sure that competing congestion control algorithms operate closer to the picture on the left rather than the picture on the right, which is what happens in a classic FIFO queue. More more specifically, what we argue for in the paper is that most congestion control algorithms, because they make these trade offs that exist on the Pareto frontier, what we wanna minimize is this movement, and we call this minimization of movement performance isolation, where we say that they're able to meet their performance goals as if they have been isolated by the other flows that they're competing with. Now the obvious solution for performance isolation or making sure these flows don't move around is. We can give each of the flows their own queue, and this way they're able to fill the buffer as much as they like, and then everyone's happy. Everyone's making the trade offs that they can make. But often on the Internet, you have too many, too many flows and too few queues. So, really, what this goal dilutes into is how we can provide this performance isolation with just a handful of queues. So if I have tens of thousands of flows, how am I going to isolate them from each other if I just have a handful of queues available in my switch? So in order to make this, practical goal for the Internet, in our people, we make two key contributions or two key observations, actually. The first one, is that often you don't need perfect isolation between these flows. You could make do with approximate isolation between these flows. So because a lot of congestion control algorithms exist on the Spirito frontier, you could actually cluster them based on how many queues you have available, and that way you could provide what we call approximate performance isolation between them. So for example, over here, if I've got all of these flows and I've got only three queues available to provide some kind of isolation, I could cluster them based on the trade offs that they wanna make and then assign each cluster its own queue. So that's number one, which is moving from a goal of perfect isolation to approximate performance isolation. But it's still not a solved problem because we're still left with inferring the trade off that a flow might want to make. A packet in a network is just a packet. It does not come with a label, saying that, oh, I care about throughput 60% more than I care about delay. So we it's impossible for us to map a packet to a specific point in this period of frontier or assign a cluster based on just, this information. Now, fortunately, we have a technique to deal with this problem as well, and I'm going to try to build the intuition towards the solution step by step. So to start with, let's say somehow my network has converged to a happy place where all of my flows are somehow, you know, in their correct clusters and in their correct queues, and I have a correct assignment. And in this correct stable assignment, I have a new flow that enters, and I'm gonna color this flow black because I don't know where it exists on the Pareto frontier. Now in the state where I have no knowledge about what kind of a trade off this flow might need to make, I can start by assigning this flow just a random queue. So let's say I put this flow in queue number two, and then what I can do is I can observe how this new flow interacts with the other flows within queue number two. So if this new flow turns out to be super aggressive, that's usually a signal that this new flow cares about throughput more than the other flows within queue number two, which I can use as a signal to promote this flow up to queue number one where I'm housing all of my other throughput hungry flows. And this simple shuffling mechanism can be a really nice natural way to automatically distill the flows within your network according to the throughput delay trade offs that they want to make. And we can assign them the correct queues over subsequent shuffling rounds. So we built, a multi queue AQM based on this simple intuition where every round, we look at flows in each queue, and all of the nice flows get shuffled down to delay sensitive queues while all the naughty flows, they get shuffled up to more throughput hungry throughput hungry queue. And the reason I'm I'm calling these flows naughty and nice is because we call our system Santa because every round, it's making a list of naughty and nice flows and then shuffling them up and down. So we implemented Santa on a real p four switch. And in our real implementation, we also made a number of more practical considerations. So back for example, the first practical consideration we made was to have a dedicated MICE queue where we would house all of our short flows that would be too short to respond to any of these, shuffling behaviors. We also, gave each queue bandwidth proportional to the number of flows that it housed in order to, you know, match, flow level fairness as much as possible. But, of course, this is a configurable parameter within SANTA. We track the throughput shares of, of each of the flows using the sojourn time of each of the packets, which was then, exported up to the control plane, which would use this to make its shuffling decisions, update the bandwidth weights of each of the queues, as well as the initial assignment of flows to particular queues. And I I don't need to get into the implementation details, but all of these bits and pieces very naturally fit into the data plane and the control plane pipelines in a p four switch. So we built Santa over this, you know, intuitive idea of shuffling flows according to the trade offs that they wanna make. Now let's see if it actually does meet the goals of approximate performance isolation. So let's just set up in the paper, we have many experiments where we try to convince the reader that this does indeed happen. I'm gonna go over one of these critical experiments, in the talk today. More specifically, I'm gonna take a very simple scenario where we had nine long running flows, three each of Cubic, BDR, and Vagus, and we configured Santa to shuffle them every five seconds across a system of three queues. And, to set the context, this is our baseline, which is a simple FIFO baseline, where if you just put all of these nine flows in a single queue, you can see that cubic and BBR take most of the bandwidth, at least in this buffer network setting, while Vegas gets hardly any of the bandwidth. And what we're aiming for is a fair queuing system where we ran the same experiment but assigning each flow to its own queue. So what we deal with in most of the Internet today is on the left. What we're targeting is the graph on the right. So how did we do? So it turns out Santa does really well over long enough shuffled intervals by isolating flows in the individual queues and clustering them correctly to the, to their correct clusters. And we can sort of observe this in Santa's graph by seeing that there's very little translation from where these flows operated when they were run-in isolation or when they were run-in fair queuing compared to where and what kind of trade off they end up making while they're passing through a Santa like multi queuing system. In the paper, we compare SATA with many other approximate fair queuing systems that that unfortunately are not able to give these approximate performance isolation guarantees. Now, given this this performance, I think another natural question that we should discuss, before we end today is how many queues is enough for Santa. Right? Because we're naturally talking about multiple flows that can organize themselves in multiple clusters. The natural question is how many queues you need to provide this amount of performance isolation. So we provide analysis for this in the paper as well. And what we show is that Santa actually forms this really neat continuum between FIFO system that has absolutely no performance isolation compared to a fair queuing system that has perfect performance isolation. So what this means is that if I ran Santa with just one queue, it would run exactly like a FIFO system. But as I start supplying Santa more and more queues, we can see that the operating points of the competing flows sort of start meandering towards their ideal, operating points. And indeed, when you supply Santa with enough queues, which is the number of flows, Santa behaves exactly like a fair queuing system over time. In the paper, we have a lot of scalability experiments as well where we show that, you know, these performance isolation guarantees can be met for a lot larger number of flows as well. Naturally, the more number of queues you have, the better performance isolation you can get. But what we have seen with Santa is that at least over long long enough time horizons, with just a handful of, number of queues, you can, often manage to get significant approximate performance isolation. Now I've presented all the nice results from the paper, but this is probably a good time to point out that SATA is in no way a silver bullet solution, for the Internet. And, I can continue to do congestion control research for many, many more years. And the reason I say this is because, you know, given Santa's simple design, the shuffle frequency is not an easy parameter to, to tune for a given network. So in the in the people, we have mainly looked at long running flows, but office obviously, you can imagine the shuffle frequency is going to be a tricky parameter to tune depending on what kind of flow churn you have within your network. We also need to explore more dynamic buffer and bandwidth allocation strategies. So maybe it would make sense to give the throughput hungry flows more of your shared buffer compared to your delay sensitive flows, and the same thing goes for bandwidth as well. Finally, this is a new queuing, mechanism. It would also make sense to analyze this game theoretically, see if there's a game theoretic way to make sure that there's no way to game the system, or if there is, how can we modify the algorithm to make, Santa immune to this? So that that's all I have for today. I think the key three takeaways that I hope you take from this paper is that, bandwidth fairness is probably not the right goal for the Internet anymore. Instead, we should look at systems that are looking at things that are similar to approximate performance isolation. And with Santa, we indeed show that this is something that can be practically achieved even at the scale of the Internet. Finally, our Santa implementation is open source and available on GitHub. Before I end, this was work that I did with with many wonderful collaborators. And please reach out to me if you think SATA makes sense. I would love to hear if it makes sense to deploy within your networks. And more interestingly, I would be very interested to hear where in your network you would think you could deploy Santa. Yep. Thank you, and I'll be very happy to take questions. [01:40:44] **Simone Ferlin**: Hold on. Mic. [01:40:47] **Jana Iyengar**: Am I up first? [01:40:48] **Michael Tuexen**: Yes. Oh, okay. Excellent. [01:40:50] **Jana Iyengar**: Thank you. Jenna, Iyengar, thank you, Ayush, for presenting this. I think this is fantastic work. I read this as basically, you know, we've had f q in the Linux kernel, for example, but that does require per flow state. And I see this as sort of a next level up in terms of congestion points for providing flow isolation where we can't do per flow state. So that read is right. You're nodding. So I'm assuming that is the correct read of this work broadly. [01:41:18] **Ayush Mishra**: So for so the the main read is not just per flow state, but also the number of queues that you might have a read. SANTET does require per flow state because it does still [01:41:34] **Jana Iyengar**: Yeah. I get that. The shuffling and everything else will require it. Yeah. [01:41:37] **Ayush Mishra**: But it does apply this to only a subset of the flows. So for example, if the flows are short, you just put them in the MiceQ, and that's how we manage state. [01:41:44] **Jana Iyengar**: Got [01:41:45] **Ayush Mishra**: it. But if all of your flows needed, it still does not solve the workflow state problem. [01:41:50] **Jana Iyengar**: But the number of flows that you're showing there are pretty reasonable. Yep. So I would say that this is fantastic because of two reasons. One, you're challenging or you're at least articulating the premise that we've, as a community, known for a long time, that is that flow isolation is something we really, really, really want, and we keep testing things without flow isolation in the network, and you're doomed to failure when you get there. So I appreciate that point. I think that's very, very important. My question is how far do you need to go beyond this, beyond FQ, beyond Santa to actually, like, you know let me let me have you looked at l four s, for example? Yep. And have you do you have any comparisons and analysis of where you might be able to do with approximate isolation versus l four s style more specific isolation? [01:42:39] **Ayush Mishra**: So I would imagine l four s would definitely do better, especially in a case where, let's say, Santa also has only two queues available to it. The main place where I think Santa differs from l four s is that it does not explicitly require any cooperation from the end host. So in l four s, maybe a client could lie about what its preferences are and then get into the better queue. But with Santa, that's not gonna be a problem because I'm gonna directly observe what your sending behavior is, and then that's gonna translate into what kind of a trade off I allow you to have. [01:43:14] **Jana Iyengar**: It's a strong benefit. I hope you point that out on the paper. [01:43:17] **Michael Welzl**: Yep. Thank you. [01:43:21] **Participant**: Well done, Thanks. This is really interesting. I think one of the key issues is that you have also a queue for Mice flows because short flows would suffer actually from reordering. So I didn't read the paper actually, but so how long are the flows that you put into your shuffling queues? [01:43:46] **Ayush Mishra**: So how we classify MICE flows are basically the first flight of packets for any flow. So the first 10 or 20 packets of a flow are always put in the MICE queue, which is a strictly high priority queue. And then everything upwards from that gets assigned a queue within the Santa queues and gets shuffled around. [01:44:04] **Participant**: Yeah. So I I guess the let's say, the the the disadvantage of shuffling the packet packets is probably getting reordering, which maybe may lengthen your flow completion time. But if the flow is quite long running, then that is basically another problem. So did you look at slowdown metric? [01:44:27] **Ayush Mishra**: So we looked we did not look at slowdown. We our primary metric was still how much it translated in terms of what the average throughput delayed trade offs were. But we did look at reordering, which was pretty insignificant. In fact, in most of the experiments that we ran, it did not happen. Mhmm. And the reason we'd I think this happened was because we very carefully made decisions on, you know, whether we shuffle the the nice flows down first or the naughty flows go up first and when in which order we dequeue from each of these queues. I think that sort of manages the the reordering behavior of it. [01:45:04] **Participant**: One quick last question. Did you compare it to flow queuing? [01:45:09] **Ayush Mishra**: I don't believe we did. Okay. But I I can chat with you [01:45:13] **Participant**: Yeah. [01:45:13] **Ayush Mishra**: Offline regarding this. Yeah. [01:45:17] **Stuart Cheshire**: I'm Stuart Cheshire from Apple. I think it's a sign that you're doing good work and giving a good presentation when there are this many people in the queue to make comments and ask questions. So congratulations for that. [01:45:33] **Ayush Mishra**: I I'd very happily take that as well. [01:45:36] **Stuart Cheshire**: Yes. You can benefit from all kinds of feedback. That's true. But it is better if it's delivered politely. So good work. Very impressed. I liked your point that you made that too many people assume that throughput is all that matters. So I'm happy to see you making that point. That's absolutely right. In the days of dial up modems, we all wanted more throughput. But now that you can get 10 gigabit service to your house in some places, it's not all about more throughput. So good point there. One thing you didn't say, but it was I felt it was kind of implied, not just in your presentation, but in other presentations that it keeps coming up. There's an assumption, an unconscious assumption we fall into that there's an either or choice, that you can have high throughput or you can have low latency, and you have to pick one. Yep. And and the insight of the Kleinrock point is this this unconscious assumption we have that there's some linear relationship that as throughput goes up, latency goes up with it, and you have to decide where to be on that spectrum. That's not true. Throughput can go up with the latency being unchanged Yep. Until you fill the BDP. And then beyond that point, the throughput doesn't go up, but the latency starts going up. So I think there's a really key insight there that you can actually have the best of both worlds. And that sounds too good to be true, which is why people forget that that's possible. [01:47:04] **Ayush Mishra**: Yep. [01:47:07] **Stuart Cheshire**: So now I'll get to my questions and my suggestions. All of this analysis and, again, this doesn't just apply to your presentation today. There was no mention of ECN, L4S, the Prague congestion controller. And I think in the modern world, that has to be compared. And in some ways, might be disappointing because if it solves all the problems, then there's less research to be done. But this is not theoretical at this point. Right? Your your your iPhone and your Mac have have TCP Prague, and Comcast has deployed L4S marking at the bottleneck. So this is transitioning out of the lab. And you said something there about traffic lying to get in the better queue. And I think that's a common misunderstanding about L4S. We've been doing traffic priorities and type of service and traffic classes for thirty years, and it hasn't solved the problem. L4S is not just a new name for the old bad idea that doesn't work. It turns it on the head. L four s is not a better queue. L four s is better information. By marking your traffic ECT one, you're not saying give me special service. What you're saying is give me better information about the state of the queue so that I can make more intelligent decisions. So there really is no incentive to cheat with L4S because the winner of L4S is using Prague or something like that that makes smarter decisions. If you don't have your send and make smarter decisions, the fact that you promised the network you'd be smart and then you're not smart, you're shooting yourself in the foot. So I guess I've talked enough. We have other people waiting. But, overall, good work. Thank you. [01:48:59] **Ayush Mishra**: Yep. Thank you so much for the feedback. I I would probably loop back to the first thing that you said, which is, you know, it's a trade off between throughput and delay. I agree with that, except I think the trade off space goes beyond two dimensions. So, you know, depending on if you wanna optimize for loss, jitter, if you have short flows where you're optimizing for FCTs, then, you know, even though you can make that trade off, it might make sense to access somewhere further down the line because it's optimizing some other dimension. [01:49:32] **Chris Box**: Sorry. I'm Chris Box. I work for BT. Thank you for the presentation. I find it really interesting. And also the slides are nice and clear as well. I appreciate that. The so this this is the system is something you're comparing to FQ and you're saying FQ gives you ideal isolation between flows. And this is this is a way of doing it with a small number of queues. So so, obviously, I'm I'm I'm just thinking about what does this mean for resource requirements. And clearly, you need less memory if you're doing Santa. What do have you got any measurements on what it means for CPU? Are you do are you using more CPU because you're doing the shuffling or using less CPU because you have fewer queues to manage? Do you do you have any numbers? [01:50:25] **Ayush Mishra**: So we haven't implemented this on the endpoint on the kernel. We have implemented this in a programmable switch. So over there, of course, the shuffling behavior does take some CPU cycles from the control plane, but these are cycles that are available unless you're running many other algorithms on the switch itself. With regards to your comment on memory, I'm not sure Santa would require less memory. It does require a few queues, but I think the amount of buffer storage and the memory that you wanna use is going to depend more on what your flow churn is, what your BDPs are rather than a system like Santa. So those are things that I don't think Santa can make up for, but it can allow you to, you know, edge closer towards this goal of isolation while having fewer queues. [01:51:14] **Gorry Fairhurst**: Okay. [01:51:14] **Chris Box**: Yep. Yeah. Makes sense. Because my my mental model is is just like a Linux kernel being the router. And so everything's in everything's running in in the kernel in doing yeah. With normal CPU. So I get and I clearly yeah. I understand that with with different hardware, that trade off is different. [01:51:37] **Ayush Mishra**: Yeah. [01:51:38] **Chris Box**: Alright. Thank you. [01:51:39] **Minqing Wu**: Thank you. [01:51:41] **Participant**: Hi. Danish from MPI. I have a few questions, but I will start with with with the Symbols one. Have you tried this in the multiple bottleneck? Because in the first bottleneck, you are already touching the dynamics of the flow. So in the second bottleneck, it might get impacted. I mean, the the objective of of the flow is changing. How do you set set the objective for the next one? [01:52:05] **Ayush Mishra**: Yeah. That that's a great question. That's not something that we have looked into in a lot more detail. I think it's it's gonna be very interesting to see because depending on where which queue your flow gets shuffled into, can potentially have moving bottlenecks where you should get shuffled into a queue where this is no more the bottleneck anymore. Now your bottleneck moves somewhere else. So you could have shuffling bottlenecks within your network as well. So I think it's a great question. It's a very interesting dimension, but we haven't investigated this yet. This is something we're still looking into. [01:52:38] **Participant**: Okay. I would I have one more. About the flows with different RTTs because RTT is changing the objective. I mean, objectives are the same, but but but the RTT make might make a flow more throughput hungry than the other one. Right? Then how do you set the fair queuing for these? [01:53:01] **Ayush Mishra**: So I I think that's the beauty of Santa, which is that it does not care about what your RTT is or what you it's gonna purely behave based on what you're sending behaviors. So even if I had all nine flows being cubic, for example, and they had different RTTs, they would be shuffled across, according to the aggression that's a function of the RTT rather than the algorithm. So, yeah, to answer your question, it could still help with RTT and fairness as well, and it does not explicitly try to do this. It's just the natural sharpening behavior achieves this. [01:53:37] **Gorry Fairhurst**: Thank you. Nice [01:53:38] **Minqing Wu**: talk. Thank you. [01:53:43] **Gorry Fairhurst**: We have a mic. Hi. I'm Gori Furhus. First of all, thanks for bringing this here. And I was gonna ask about RTT on fairness, but you've talked about that. I'd encourage you to perhaps look at more queues than just one or two, but probably a handful is probably enough. And I think for me, that's the beauty of what you presented, the the simplicity. Yep. So I'd love to see a comparison sometime. Maybe you've already have thoughts against L4S in terms of the complexity of actually implementing this. Have you thought about that yet? [01:54:20] **Ayush Mishra**: So we did look at so it's not a very complex multi key Kilometers to to implement. Mhmm. And we sort of try to demonstrate this with the practical implementation that we have in p four. But, of course, we have not compared this with l four s. So that's something we would have to look into. [01:54:40] **Gorry Fairhurst**: Okay. And finally, when we wrote RFC nine seven four three, which you may have come across or not, we talked about harm. I know and rather than taking bandwidth or some other metric, look at how much you harm other flows. I think I understand that's kind of what you're looking at, the the impact on the other flows rather than the benefit of that particular flow. [01:55:00] **Ayush Mishra**: Yes. So so I think there are multiple ways to approach the issue of harm. If you just look at the shuffling behavior, it's trying to minimize the bandwidth harm that it's doing. But from a more theoretical standpoint, if you isolate flows, then it makes them impossible to inflict harm. So that's also one way to minimize harm, which is if you know flows have, shared goals, then you can club them together and that should also, in theory, minimize overall harm, whatever your harm metric might be. [01:55:29] **Gorry Fairhurst**: Thanks for your talk. Thank you. [01:55:34] **Mohit P. Tahiliani**: Mohit, hi, Ash. Thanks for bringing this here. I have one question and one comment. The question is I I was just looking at the paper. Have we tested with 90 flows maximum, or is it any number beyond 90 that we have tested? [01:55:52] **Ayush Mishra**: We have only tested with 90. With 90 flows. Yep. Okay. But these are, like, real flows from the Linux call. [01:55:59] **Mohit P. Tahiliani**: Oh, yeah. I saw 30 cubic, 30 b b r, 30 Vegas Yep. Across three queues and then [01:56:05] **Chris Box**: six queues. [01:56:05] **Mohit P. Tahiliani**: I was just looking at it. Thanks. Okay. The comment is about the the word FQ. So I think we are I I see it's mainly referred as FAIRQ. Yeah. There's also FQ, which is slightly different, and there was a little bit of similarity that I find, and maybe it will be very interesting to see if we could see how Santa can further be combined with FQ also because FQ has a property of maintaining two different sets of queues. Of course, there are many different queues, but two sets of queues. One is new flows and one is the old flows. Yep. Beyond that, it has something called sparse flow prioritization. So your mice flows automatically get higher priority during the DQ using a slightly modified version of DRR. [01:56:50] **Minqing Wu**: Yep. [01:56:51] **Mohit P. Tahiliani**: And I believe that we may be able to reduce the shuffling by a lot if you try to incorporate and see if it is a flow kicker. FlowQ naturally, does this property of you know, it has a property of prioritizing this fast flow for the [01:57:05] **Ayush Mishra**: market. So, I think we we came across many similar ideas and, you know, other data center flow management systems as well where, you know, you treat the first few packets. I mean, obviously, you don't know how long a flow is. So each new flow sort of gets treated as a mice flow, and it gets its own MICE queue. And that is something that we try to implement in Santa as well. So Santa also has a it has a single dedicated MICE queue where the first 10 packets of each flow are shuffled with high priority. [01:57:35] **Mohit P. Tahiliani**: With FQ, you will have several queues for MICE Yep. Not just one queue. Yep. And then then again, there's a further priority on that. So maybe it will be just interesting to see if Santa can be on that. [01:57:47] **Ayush Mishra**: Definitely. Yep. Alright. Thanks. [01:57:52] **Michael Welzl**: I want to say almost exactly the same thing as Mohit, maybe with a little less knowledge about the same thing that he's talking about. So the primary thing for me to say is that I'm excited about it. I think this is really, really good work. That other thing so what I was so so maybe I'm misrepresenting the FQ behavior. My theory has been not knowing the exactly that a new flow in that system that is prioritizing new flows may be a flow that just doesn't have anything in queue. So I think we may have a case where there is a flow that just doesn't have a queue. And you got a new packet from that flow. You basically treat it almost instantly, and it goes away, there is no queue. Right? Whereas in your system, if I'm imagining a flow that you know, a sparse flow, like a voice over IP flow or a gaming flow or anything that isn't really queue building, it might end up being queued together in, you know, with flows that that do have a queue. Right? So we do have [01:59:00] **Ayush Mishra**: a time out on how long the flow must be active. So what we do is we have, like, a we have a table where we index each of the, four tuples that we have seen in the past fixed horizon. And within this fixed time horizon, if you send less than 10 packets, you're always gonna be within the MICE queue. If you send more than 10 packets, then you're gonna be promoted within the Santa queues. But, I mean, you you it is a very important aspect that you that you pointed out at. You know, the tuning these parameters is gonna make a massive difference in terms of, you know, how many things end up in the MYSEQ, how many things end up, and how often you shuffle them as well. So, yeah, I I completely agree. These are parameters that we don't currently have a good strategy for tuning, but this is this is something that we would have to look at. [01:59:49] **Michael Welzl**: I I also think, like, there will be ways to tune that more and more and more, but I I think this is the absolutely right direction to go. So that's really great. [01:59:57] **Jana Iyengar**: Yeah. [01:59:57] **Ayush Mishra**: Thank you. [02:00:01] **Jana Iyengar**: Thank you. Just on the FQ question, when I mentioned FQ in the Linux kernel, I did mean [02:00:07] **Minqing Wu**: Flow q. Flow q. [02:00:08] **Jana Iyengar**: Okay. Okay. Bob Briscoe is not here in person, but his spirit hangs around in these hallways, so I'll invoke it. Have you looked at how applications application utility tends to be application utility flow is thing that we like to engage with, but applications don't care. And then in a queue system, typically, you will have this problem that gaming can it can be gamed. Yep. Have you thought about it? Have you done any simulations to figure out how that performs with multiple they get multiple bandwidth shares is really what I'm asking for. [02:00:39] **Ayush Mishra**: We haven't looked into it yet. [02:00:40] **Jana Iyengar**: Yeah. I would say that's an important thing to consider. Yeah. And you mentioned that you had to do some game theoretic analysis. This is actually fairly simple in terms of just doing a simulation and seeing what [02:00:48] **Michael Welzl**: actually happens. [02:00:49] **Ayush Mishra**: We we would still have to yeah. I would still have to look into it. I don't have a good answer for you right now. So I I agree. Flow level fairness has has the has the No meaning. [02:01:01] **Michael Welzl**: Yeah. It has That was [02:01:03] **Jana Iyengar**: a that's the face you're looking for. [02:01:04] **Ayush Mishra**: Yeah. Thank you. Thank you. [02:01:11] **Simone Ferlin**: I think that's it for today. Thanks for being with us for two hours. And yeah. See you next time.