Session Date/Time: 20 Jul 2026 09:30
[00:00:22] Georg Kot: Okay. Welcome, everybody, to the second slot of this year's ANRW. My name is Georg Kot. I'm sharing this one for for you. And we have five security talks, three long and one short one, starting out with Paul Schmidt presenting unequal cover measuring privacy disparities in proxy infrastructure. And be come in, find your seats. And because we are short on time, we have a density pack program. I suggest that we hand over the floor to him, Paul. Thank
[00:00:58] Tommy Pauly: you.
[00:00:58] Paul Schmidt: Are you able to hear me?
[00:01:01] Suresh Krishnan: Yep.
[00:01:02] Paul Schmidt: Okay. Great. So I'm Paul Schmidt. I'm an assistant professor at Cal Poly, and this is work with three truly excellent undergrads that I worked with over the last year. We're looking at privacy disparities. So just as a little bit of background, what we're studying here are multi party relays. You likely have some idea about these, but if you don't, probably the most popular is Apple's iCloud Private Relay, but there are others. And the main idea behind these systems is they split trust among non colluding non colluding intermediaries. So Apple runs the Ingress server. They're going to see your identity in the form of your IP address or some sort of token or or something like that.
[00:01:49] Suresh Krishnan: Audio is clipping. Can you hold the the line mic closer to your mouth, like, and hold it while you speak? Because when you move, the audio is clipping in the room here. So just get the whole mic closer to your mouth on the on your headset on your head on your headphones.
[00:02:06] Paul Schmidt: It's it's it's the mic in my camera, so I'll try to I'll try to not move so much.
[00:02:14] Georg Kot: Yeah.
[00:02:15] Lohit Karyala: Okay.
[00:02:16] Paul Schmidt: So the the idea is they split trust across parties. So Apple sees your identity, and then their their partners, in their case, it's Akamai, Cloudflare, and Fastly. They see your destination but not your identity. And this works. This gives you architectural privacy, but what we're curious about with this work is does this give users equal privacy across the board? And so we're studying anonymity, and and we're we're studying many different metrics about this. There are lots of different properties that these sorts of proxy systems provide. They give you IP hiding. They give you flow unlinkability, etcetera. And the main idea is that these egress IP prefixes are shared, so users are hiding in a crowd. And with any privacy system, larger the crowd, the more anonymity you're granted. So we could go to one logical extreme. There's a single IP prefix that all users are egressing onto the global Internet, but the downside of that is geolocation. So they assign these IP prefixes and then tell the world that this is located in San Jose or New York or London. And if we all came out of one egress location, that would obviously be a problem in terms of streaming video or, you know, trying to go to homedepot.com. And so in response, these operators tend to map users to nearby egresses. And so what this work is doing is measuring geolocation anonymity, where we're looking at the set of users that are sharing a location and egressing on the Internet, and we're studying whether commercial NPR deployments provide equitable geolocation. Like I said, this is a case study about private relay, although I will say this is not private relay specific. The reason we study that system is because it's measurable where others are not. They publish data. They they make it available. So they they release this public dataset with all of their IP prefixes and the locations that that users are supposed to be associated with. And that gives us the egress location information for studying infrastructure. We focus on The United States in this study, and so we look at all of the egress locations in The US and then cluster them based on locations. So if they're overlapping, we essentially just bin them together if they're within one kilometer of each other. And then for users, when we're modeling usage, we use census tract data. So there are 83,600 US census tracts. And the nice thing about that data is it gives us other information like income or education, etcetera. And we use a baseline of mapping to the nearest in state egress for any census tract, And the paper goes into some of the details. We do some testing to make sure that this actually makes sense, which I'll talk about in a bit a minute. So in terms of privacy metrics that we're looking at, first is crowd size. It's very basic. That is the sort of maximum number of users that that you could be hidden within in a location in an egress. And then we look at degree of anonymity, which gives you a normalized value. So it doesn't depend on the the size of the crowd anymore, but it's dependent on the adoption rate. So we use a flat adoption rate that was picked up from APNIC, posted that something like point 11% of of users are using Private Relay. And then using that degree of anonymity, we can calculate inequality. So in a region, so, say, a state or a county, we can look at the tenth percentile and the ninetieth percentile degree of anonymity and compare them. And if you see very large differences, now you have inequality. We can reason about that. And so the first result is is very what very much what you would expect. Density is going to predict your crowd size. This is not all that shocking or interesting, but we do think the magnitude is is something to at least note. So we we cut or we slice the tracks up based on rural or suburban or urban definitions, and we see that rural tracks tend to have 9.3 x smaller median crowd sizes compared to urban regions. Of course, this this narrows at the tails, but but it's a significant difference. So we're seeing 39,000 median versus 370,000 median crowds. So underneath that, we were trying to understand what's actually driving this or or what's what's influencing some of these factors, and it turns out income is is actually quite a powerful thing. So in these rural areas, what we're plotting here on the left is median income of rural census tracts versus their crowd size, and then on the right, we have urban. And you can see that in rural tracts with high median income, you actually end up with a larger crowd size on the whole versus the urban, it's basically flat. We don't really see a strong correlation there. So what that means is in rural tracks, you've already got a lot lower crowds or smaller crowds, and you're income sensitive. And so you have this compounding disadvantage for users in these areas. We study this with other metrics. So so education shows a very similar pattern. And what we take away from this is that urban density just provides a baseline of privacy that rural regions can't overcome. So we go into this a little bit in the paper, but we wanna make sure that we weren't just selecting Our our our mapping of users to egress wasn't actually sort of changing the numbers all that much, and it turns out this holds across k two, k three, etcetera. The one thing we did test that was kind of interesting is income weighted adoption. So, like, I I call private real estate paid service. And so if we don't go with a uniform adoption but instead actually have income weighted, it amplifies inequality. So we think that the the results in this paper might actually be sort of conservative, and and things might be much worse in systems where people are paying money. And so what we were trying to understand is what is causing inequality that we were seeing. And so here we have the tenth percentile. So we took all the counties that had at least 10 census tracts, and we look at the tenth percentile degree of anonymity versus inequality and then the ninetieth percentile degree of anonymity versus quality. And you can see the tenth percentile is what seems to be sort of correlating with the inequality. We have a a negative correlation of negative point six nine for tenth percentile, but the ceiling is is not all that that strong. So our takeaway here is that what we're seeing is the problem tends to be under protection of the worst served, not over protection of the best served in terms of driving this inequality where we see it in these locations. And so the nice thing about that is that means we should be able to potentially do something about this. And so we run a little simulation, and our idea here is we didn't want to introduce something very egregious in terms of trying to change the system. We're not trying to argue for additional new infrastructure in places. And so instead, we're looking at taking away egress locations. And so we're iteratively removing them, and then we're reassigning those tracks populations to the next nearest egress and recalculating all of these these metrics. So we did this across four different algorithms. We did smallest first, so the egress serving the smallest population. Remove that. And we did equity greedy, which is trying to lift that tenth percentile as much as possible. We did distance greedy because the downside to removing egresses is now you're going to move users further from the egress that they're assigned to. So when I go to homedepot.com, I wanna be able to get to you alright. But I want my egress to look like I'm actually showing up in San Luis Obispo, California, not San Jose. And so there's this penalty. It's not free. And then we also looked at random, so just removing egresses. And so what we see is we remove about 19%, and that grows the crowd size from 5,800 to 14,000, a 2.5 x increase. But there is this downside. Use this was equity greedy. We see a mean tracked egress distance grow from 8.6 kilometers to 10.2. The interesting thing here is that smallest first actually works quite well. You don't need a really complicated algorithm. It it nearly matches that equity free approach. And so just to wrap up, these this proxy placement really inherits and amplifies existing digital divides. This is not unique to private relay. This is not particularly a study about private relay. It's a study about systems like this. We see these sorts of systems following infrastructure and population geography. But we do think that inequality can be remediable without building new infrastructure. Of course, there's the sort of danger of overloading these systems. We have to look at usage to be able to remove egresses in this way. But we believe that as these sorts of systems become more and more commonplace, minimum crowd size targets might actually be a metric that we should be paying attention to rather than simply focusing on performance. So there are several limitations with this system. We don't know the actual Apple mapping from user to egress. That's proprietary. We try to mitigate this with sensitivity analysis, but, of course, we don't fix it. And this is cross sectional. This is this is a single system in a single point of time, but we think that this is an interesting enough result that that we wanted to share. So with that, I'd be happy to take any questions.
[00:12:41] Georg Kot: Thanks so much. We've got time for questions.
[00:12:43] Tommy Pauly: Alright. Thank you, Paul. This is Tommy Polly from Apple. We've talked before about this, but I I appreciate your digging into this. I I do think this is a a generally interesting problem to look at when you're trying to have privacy systems like this. So I had a question about, like, when you get this, how how are you measuring this? Like, is this all just based on the stats of the populations we predict in the regions? Like, when you're talking about, like, the adoption rate, is that just extrapolating the numbers, or is there any actual empirical measurement?
[00:13:17] Paul Schmidt: Yeah. So so this is pure uniform. We take the population. We assume point 11% actually actually subscribe. So, obviously, there are many flaws in that.
[00:13:29] Tommy Pauly: Got it. Yeah. And it it be this makes me interested in trying to measure some of the actual rates here. Mhmm. Mhmm. And I just also wanna note, like, as I'm sure you know, but for this group here, like, you know, there are modes, for example, for private relay where you can select, I want to be at a time zone granularity, which I think gets you towards it it seems very similar to what you're pointing to of, like, expanding the granularity here. So did you think about how that would change the anonymity sets here? Would a suggestion be to kind of have, like, a middle ground there? Is that what you're thinking of?
[00:14:03] Paul Schmidt: Yeah. So we didn't explicitly measure that, but I think that'd be really interesting to sort of look at. Is there some baseline floor where maybe we start to push users towards something other than the closest egress? Because in rural Kansas, you may be the only user using private relay, and now you're not exactly anonymous in that way. So, yeah, yeah, some sort of dynamic system like that might make sense.
[00:14:27] Tommy Pauly: Alright. Thank you. And then just, like, one last comment, which is I don't think could be answered here, but it'd be good to, think about also disparity between countries, like because I think this is just US. So regions around the world would be really fascinating to look at.
[00:14:42] Paul Schmidt: Yeah. It'd be great to have actual data, so we should talk.
[00:14:45] Suresh Krishnan: Thank you.
[00:14:49] Georg Kot: K. Thank you. Any any other questions? Looking at the time, this is also practical. Thanks again, Paul. Also for being that early. Okay. Thank you. This gets us to our second presentation today. Given if if I'm not mistaken by, Virit Shibakaschasida. You're coming up the stage. Great. She's going to talk about IoT device fingerprinting and floor's use.
[00:15:39] Suresh Krishnan: Do you wanna use the clicker? Do you wanna use the clicker?
[00:15:43] Virit Shibakaschasida: Okay. Press
[00:15:44] Stuart Cheshire: it. Yeah.
[00:15:46] Virit Shibakaschasida: So hi. Hi. My name is Verad, and I'm going today to present you our paper, adding the wired labeling IoT device type using passive network telemetry. So let me start with a problem. So IOTs are everywhere today, and the market is exploding. We are talking about thousand of different vendors and 100 of device type. And it never stands still, the everyday new devices in the market. And for someone who want to secure his network, first, he need to identify the device he have on the network so he know how to protect them. But when you try to identify the device, you get unknown device. And the unknown device is the one you can secure. And to understand why you get unknown device, I want to start with the identification method you have today. So there are two main approach. The first approach is using machine learning. So, basically, you you learn the device signature, things like IP, ports, package count. And it's work good, but in order it to to work good, you need to have a training sample of each device in the market, which in the IoT world, it's it's not feasible to have samples from every device. And the other approach is to your is to use OUI. And so, basically, the MAC address prefix map to the vendor. And the problem here is, first, it's mapped to the network chip maker and it's not always the same as the IoT vendor. And second, it don't give us the type. It don't give us the device function, just the vendor. And there is Phinig. It's a very leading product that's doing identification using crowdsourcing. And also to do a good crowdsourcing, you need to have samples of each device in the market. So this is the exact problem we are trying to answer. How can we identify a device that was never seen before? And our solution is a system that label never seen device, identified vendor and a function. So our solution called Zelle, and what make it different is first, it's work on unseen device. It's labeled device, labeled device never seen before fully offline. Second, it's zero shot. It need no training sample of the device. So, basically, each device you gave to our system, you get a label. And it's passive. It use all just the data data we can observe, and it's not doing active probing. Also, we use IoT catalog. We will see in few slides how it help us. And it give us a label. The label contains vendor, function, confidence score, and justification. So a human can verify the result and to to know if the label is good or not. So we are doing labeling, not identification. So we are not real time. We are working offline. So if the identification method give us unknown device, you just can go to our system, get a label, and then use this label to train the identification method. So let's go deeper into our solution. So the first things we'll do is the extraction. So as I mentioned, we take we take the observed data we see, and we take clue from this. And the clue can be things like domain, make, host name, and things we can just take from the network telemetry. And this is on it's on saying nothing to us. Okay? It's just string, and here comes the important part of our system. We are doing enrichment. And the enrichment is basically search the features on Google. And why it help us? It's because we want to use the the knowledge we have on the Internet. So we take into con we believe that people write user admins and blogs and think about their devices, and it can help us to label the device that we don't know. So let me give you example why the enrichment is so important. So let's say I have host name LB130. So it's on again, I don't know what it is. It say nothing to us to me. But when I search it on Google, you see I just get it's TP Link smart light bulb. So I I just get the context from searching it. So after I have the enriched data, now from the enriched data, I can just label the vendor and the function. So how do I do it? So first, I have a catalog of of all the vendors and all the types that in the market. And first, we are doing the vendor labeling, then the function labeling. And the reason we are doing first the vendor labeling is because most of the IoT vendors are specialized. So it means that about 80% of the vendors in the market only do one or two device. Let's say I know I have a Ring device. So I know it can be a doorbell or a camera, but it cannot be a printer because Ring don't do printers. So if first I know the vendor, it's it's much easier to me to know the device type. So I have two catalogs. I have a catalog of with about one thousand and forty hundred vendors and about 41 function in the catalog. So vendor labeling. So this is very simple. It's just string matching. So we just look at each vendor in our catalog, and we count how many times it appeared in the enriched data. And the vendors that get the highest score is the vendor, and also it give us the confidence score. And the reason it's so simple is because it's just about to recognize the name. We don't need to understand the meaning, and just we need to recognize the name. So other model just don't do better result. So to understand what drives the vendor labeling, we check each feature alone, and we can see from here two things. First, the enrichment is important. Let's look at domain. Before the enrichment, we succeeded about 60%, and after the enrichment, it jumped to 82%. And and to do great result, we need to combine features. So, also, we we compare the vendor labeling to other solutions. And let's see. ZIL is leading with 86%. It's about it's above GPT and. And, also, if we just took the OUI lookup, we get 64% again because the OUI give us the maker chip is it network chipmaker, and it not always is the IoT vendor. So after we know the vendor, now we are doing the function labeling. And we use and this is way harder. So we are doing or we took Roberta. It's a model, zero shoot classification, and we we don't did did the fine tuning for this. We just it gets inputs and rich data in the catalog, and it give us the label with confidence score to each of the of the candidate of the functions. Also, here, we check what drives the feature. And we see here it's way other than the vendor. Right? Here because here, we need to understand the context. We need to understand the meaning. It's not it's not just about recognizing a name. So it's way other. But if we use the function catalog, so after we know who's the vendor, we can narrow the search to just the functions this vendor make, it give us better results. And better results is about 25 improvement if we don't use the catalog. Also, here, we compare our result, and we also see that our product is leading with 75% above GPT with 64%, but 62%. And, also, we took a random baseline, which basically we choose randomly from the catalog, and we got just 40%. So our solution is working and in working better than random. So so here we see our result. We test our system on two different datasets. First, on a dataset from the lab, about 94 devices from the lab, also a big cyber security company give us data from real customers, about 1,122. And, also, here, we can see we got 86% vendor labeling. And for function labeling, we got 55% on the lab data and eighty eighty one on the wild dataset. And this is for each one. We also check each two. Each two mean the second best result because our solution is meant to to be a human in the loop that can verify the label. And here, you can see we give also explainability, as I mentioned in the beginning. We give the confidence score. And the confidence score also can help us to flag if the result can be not good. Because if the confidence is score, it can maybe give us little insight. Maybe the label is not good. And also, we get a justification, and it just takes snip from the enriched data that's to understand to understand why you choose the label. So we deploy our system. You can just send the feature, and we give you a label. You can just use our online system. So to conclude, I show you our systems that can label unseen device fully offline. So thank you very much. And if you want to learn more about the research, you can just enter the link here.
[00:27:48] Georg Kot: Thank you very much. Since you hold the other mic, I'm going to use this one. Do we have questions for Vivek? I'm so we have we have one.
[00:27:59] Suresh Krishnan: Yeah. If nobody has a question, I have one. Are you aware of the manufacturer usage descriptions, the MUD specification that IETF did? So there is MUD. It's called MUD, manufacturer usage descriptions. Have you thought about, like, combining that stuff into yours if somebody provides it to use that info to, like, improve your accuracy?
[00:28:21] Virit Shibakaschasida: Again, and back to MUD again?
[00:28:22] Suresh Krishnan: MUD, like, talks about, like, for IoT devices, it says, like, it tells the network, like, what the intended usage of the device is. So I think it could be a useful input for you to consider that into your fingerprinting of that stuff.
[00:28:36] Jayshree: Actually, I have a I have a slide for this.
[00:28:38] Virit Shibakaschasida: Oh, okay. I have a slide about it. So if maybe I can show it. I can show it. Sure. It's in IDE, but maybe it will be I put it in IDE.
[00:28:49] Suresh Krishnan: IDE. Okay. Cool.
[00:28:50] Virit Shibakaschasida: I'm not sure, so it will be easier.
[00:28:53] Suresh Krishnan: Let me see if I can pull it up, but you can talk through it.
[00:28:56] Virit Shibakaschasida: Mhmm. Yeah. Okay. You ask again? Question about
[00:29:01] Suresh Krishnan: Can you improve your accuracy by using that?
[00:29:04] Virit Shibakaschasida: Yes. So it's I I think it is a slide. I prefer to show it from the slide.
[00:29:09] Suresh Krishnan: Okay. No worries. Yeah. That's fine. We can I
[00:29:11] Virit Shibakaschasida: I don't really remember the exact
[00:29:12] Suresh Krishnan: No problem? That's fine. If you can talk offline afterwards, I think that'll be good. Thank you. Thanks.
[00:29:17] Lohit Karyala: Stuart? Another question.
[00:29:20] Stuart Cheshire: I'm Stuart Cheshire from Apple. Very interesting presentation. Thank you. I'm curious to know what you think future directions will be in this space. And what I mean by that is we have attention that for privacy reasons, we wanna make it difficult to identify devices. We we just had the previous presentation about iCloud private relay and not being able to track people's activities.
[00:29:46] Virit Shibakaschasida: Mhmm.
[00:29:47] Stuart Cheshire: And then the tension on the other side is that people naturally say, but I want to know what's on my network. And they don't want other people to know what's on their network. And sometimes you can't meet both those goals at the same time. Do you see manufacturers getting better at encrypting their traffic and anonymizing their devices? How do you see this evolving?
[00:30:13] Virit Shibakaschasida: Okay. So about the our direction. So about encryption. So, yes, it's a big problem. I think the world is going to go to do more and more encryption. But our solution is behind firewalls. And we assume that the firewall is doing the description, so it'd be easier to us. And also, we want to research devices behind that also can be different difficult to take the data from it because they don't have the IoT the IP network. And because they're behind hubs, that can be very confusing. So we want to research this in the future.
[00:30:58] Suresh Krishnan: Thank you.
[00:31:03] Georg Kot: Okay, thank you very much again. Switching locations. This gets to us to our third talk on source address validation. If I'm mistaken, to be given by Libin Yu. There he is. And floor is yours.
[00:31:42] Li Bin Liu: Hello, everyone. I'm Li Bin Liu from Zhongguan Chen Lab. Today, I will present our work works, SAV-A, algebraic framework for source object validation. This joint work with Li Chen, Wen, Zhou Yang, and Dan Li. For oral introduction, a have been applied as the operational mechanisms a mechanism with specific data structures such as FIB or RIP. So in our work, we show that sub can be unified algorithmic framework with different parameters. So if you see once you see that, you can get the convergence proof, correctness guarantee, and analysis tools that doesn't that didn't exist before. So I IP spoofing enables the reflection and the amplification of denial distributed denial of service, DNS cache positioning, and the TCP session hijacking. However, despite decades of standardization, defense gap persists across the Internet. Based on the data from CADA Spofer project, one third of tested autonomous systems still permit source address spoofing today. So the primary defense is source address validation. The basic idea is at every ingress interface or every router, it decides whether a package source address given where it arrived. So the sub state, it can be per interface or per prefix decision, including the valid, invalid, or unknown. Invalid means that the traffic comes at the wrong incoming interfaces. Unknown means that we don't have the information about the prefix and its legitimate incoming interface. So the hot problem is that how can we calculate their substate correctly? A quick history introduction about the South. BCP three eight introduced a manual ingress filtering, but it feels for multi homed customers. Savvy proposed perfect probing. It can calculate the accurate sub states, but it requires all the key and squares messages to get to the state, and it cannot be scalable in at an Internet scale. URPF, including strict loose IP URPF, infers the South States from the FIB or RIB, and the EF URPF adds the customer coin for the multiproider topologies. Itf.net proposed distributed protocols exchanging their specific information. However, the fundamental problem of URPF based mechanisms is improper lock and improper meet. The strict URPF drops a little more traffic when our routing is asymmetric, and the loose RPF may pass the spoof of traffic of of because any prefix will be permitted at any interface even it appears in the FIB. So the root cause is structural. OU IPF mechanisms infer their self state based on the routing tables. However, the routing tables describes a a destination reachability, which is unreliable proxy for their source legitimacy. So some that takes a different approach. The routers exchange their source specific information between between them, and the validation is decoupled from the routing tables. So the source address in analog or the BJP, because BJP announced its destination prefix is a reachable wire node. However, SoundNet announces the source prefix p can come from the can come from the source. So it also aggregated aggregated their messages by forwarding pass, which means their prefix sharing the same next hop can be bound can be bundled into a by message, which reduces the message amount from order key and squares to order and squares, which is 50 reductions in practice. So here is a open question. The third state need to propagate hop by hop. How does does this convert, or how fast that it's we can work when the routing state changes? So the the ITF working group didn't establish it. However, the distributor routing required the sovereign host routing algebra to answer these questions. SONET also require the same treatment. So SIL a fills the gap. A key observation, both routing and SIL calculate calculated their states based on the same graph, and they propagate the information along their monotone operators in the opposite semantic directions. The routing computed the best nest hop for each destination prefix. Sow calculated the ingress calculated the region source prefix for each ingress interface. So about twenty years ago, routing had its algebra theory, but this SAU didn't have to didn't have until now. SAU a has three contributions. The first one is SAU algebra. It purported the trust level signatures with aggregation and extension operators. It approved the convergence to a unique least fixed point at d round and the soundness of the relouting state. It unified or deployed some mechanisms, including strict, loose, AP, and EFPRF. They are their they are all the instantiations of SAV-A. And the third one is inter domain composition. SAV-A proposed a file level trust lattice cap which captures their binning relationships and demonstrates that validation free policies can prove the convergence and the correctness across AS boundaries. There are three compositions, self signature, aggregation, and extension. The self signature match each prefix to a task level. Zero means reject. Top level means the local originated, and their signature from the finite lattice. The aggregation accepts a prefix if any neighbor watches for it and keep the highest trust trust. And the commutative, associative, and I a potent other joint similarities of sharing their sharing their same structure behind the router selection. And the extension means what happens at the information cross link. Each link carry carries a label, a policy filter, and a trust cap. Here, for the global south state, it is one fixed point equation per node, and the unique solution is a stable global south state. So we we describe the data structures in algorithm one in our paper. So why does this converge? There are three properties. Isotonicity, left distributive activity, and absorption at bottom. Together, they make the update function monotone or a finite complete lattice. So a dropped the strict inflation because routing minimize the cost as the pass through, so it needs to it a decrease in the signature nature, but the style accumulates the trust up upward via the maximum. So convergence rests on the task case fixed point theorem rather than the strict inflation. So for the main routes, we propose theorem six, theorem seven, and theorem eight in our paper. If you are interested, you can see the detail of the introduction in our paper. So this adds a new capability. Operators can compare their configuration by fixed points and different states after topology changes and bonded convergence at DRAMs. So all the deployed mechanisms, including other various their URPF based mechanisms and a subnet, mechanism, instantiation of SAUA. A strict URPF, other best pass tree from FIB FIBURPF other feasible pass that feasible pass tag from the RIP. Loose URPF ignore the direction. It is a full graph of from the FIB. So EFPRPF adds a custom code and StunNet adds file trust levels with policy based filtering. So notice a strict API and the loose RPF share the identical parameters and and differ oh, sorry, differ only in subgraph. And for the internal, we introduce file trust layers. The four represents a local originated. Three is the customer valid validated. Two is the peer validated. One and the two are the loose and reject rejectors, respectively. We map them onto the Gao-Rexford, a valley-free model. And we use a bound labels to represent the AI's business relationships. And any value free closed work must must traverse a wider customer link so the traffic can never increase around our cycle. So we have some evaluations. First one is about a cracking correctness evaluation and the scalability evaluation. We demonstrated that over 50 thousands of entries, there are 100% correctness agreement and among the five topology families. And we demonstrated the scalabilities and show that saw a convergence can happen at around all diameters. So these are the pure Python implementation, and there are no compression or incremental maintenance. So we have also two analysis, incremental deployment, and the deep query. We demonstrate that with we show that with a 10% top level deployment, there are 9096.4% filtering. With 20% deployment, they are 98.3% filtering. And we show that removing one peer link, there are only 0.04% state changes. To conclude, to our knowledge, SAAL a is the first algebraic framework for SAAL. It proposed trials to layer our signature with aggregation and detention operators, And it show demonstrate the conversion to a unique list fixed point at d rounds. And it unifies the south landscape, and it propose a practical control plane too without any changes to their data plane or the deployed magnums. So we open sources out here. You can try it. I'm happy to take your questions. Thank you.
[00:44:16] Georg Kot: Thank you. And perfect timing. We got time for questions. Nobody cares about source address validation. This time for a change, I don't have one either. Previous ones, had to step back. But, Suresh, Thomas?
[00:44:41] Suresh Krishnan: I'm good. Thank you very much. That is a really good presentation, and thank you very much for bringing up the SAP network here. I think it's very relevant for the idea, so thanks a lot. Thank you.
[00:44:49] Lohit Karyala: Yeah. Yeah. Cool. Thank you.
[00:44:50] Li Bin Liu: And also welcome to join the SNAT session. Thank you.
[00:45:00] Georg Kot: Okay. So we move up a layer or a few more and come to DNS over QUIC. Next presentation is given by Rohit Kadiala. Quick meets reality, and this is an online talk. Can you hear us?
[00:45:24] Suresh Krishnan: Lohit, do you want Yes, sir. Slide control? You can ask for slide control. I can give it to you.
[00:45:31] Lohit Karyala: Yes.
[00:45:36] Suresh Krishnan: Yep. You have it.
[00:45:39] Georg Kot: Okay. Then floor is yours.
[00:45:54] Lohit Karyala: Hi, Ron. I'm Lohit Karyala from.
[00:45:58] Suresh Krishnan: Hello. We don't see the slides yet. Just give us a second. Are you sharing the slides, or do you want me to share the slides?
[00:46:08] Lohit Karyala: Please share the slides.
[00:46:11] Suresh Krishnan: I'll share it then so you can just give me a second. I'll give you the control of the slides Hi, Arun. So you can move the slides around. Okay? Yeah. Go ahead.
[00:46:22] Lohit Karyala: Thank you. You're welcome. Hi, Arun. I'm Lohit Karyala from. Today, I'll be presenting our initial study on quick. It's real deployment and perform performance of in Indian Access Networks. This work is a joint effort with Jayshree Sankupa.
[00:46:42] Suresh Krishnan: Lohit, your audio is coming across very poorly. Is there a different mic you can use?
[00:46:49] Lohit Karyala: Yeah. Just a minute.
[00:47:06] Suresh Krishnan: Can you see something? No. We don't hear you. Do you wanna swap? Good. No. There's one more.
[00:47:38] Marco Davitz: So Oh, no. The the last presenter is
[00:47:43] Lohit Karyala: Hello? Hello? Am I
[00:47:46] Suresh Krishnan: Yeah. You you do sound better. Go ahead.
[00:47:50] Yevgenia: Hi,
[00:47:52] Lohit Karyala: So to begin with, let me motivate this work. Encrypted DNS is rapidly shifting to quick test transport service. DNS over quick and DNS over here with you. Quick quick offers advantages such as zero zero RTK, and no out of the line blocking and improve
[00:48:16] Georg Kot: May I suggest you turn off videos so we have fewer packet losses?
[00:48:20] Lohit Karyala: Yeah. Yeah. Sure. So could offer different advantages such as faster connection set upstream or the flexion loss recovery and. However, these benefits depend on real world deployment and network support. While previous studies have primarily focused on evaluating encryption
[00:48:44] Suresh Krishnan: Lohit, your audio is, like, very poor. I think we'll probably reschedule you to an afternoon session. I think, like, we we'll move on to the next session, and we'll do some bandwidth testing with you after, if that's okay.
[00:48:59] Lohit Karyala: Okay. Thank you.
[00:49:01] Georg Kot: Thanks. Yeah. So let's let's let's have to talk a bit later and then hopefully get this worked out in the meantime. And then we move to our fifth presentation right now. We stick with DNS and speak about the illusion of DDR deployment given by we Yevgenia
[00:49:29] Suresh Krishnan: Yevgenia, you can take the clicker.
[00:49:32] Georg Kot: And Please go ahead.
[00:49:35] Yevgenia: Good morning. So the point of this very short paper was to provide a little update on the DDR deployment research I wrote about last year's blog post. Okay. So a little background. If we only have the IP address of a recursive resolver, we do not know whether it supports any form of encryption. But with DDR, what we can do is that we send a specific DNS request, and in response, we get a bunch of SCCB records with so called designated resolvers. So this is especially useful if the encrypted service runs on an alternate port number on a different IP address, or, for example, if we want to use DOH, that's how we can learn the path. One important note about those designations is that clients must not use them as is, and they need to perform some sort of validation. So for example, they can check whether the IP address of the original query resolver appears in the TLS certificate of the designated resolver. So this, for example, would not be a verifiable designation because this example IP address does not appear in the TLS certificate of 8888. So last year, I published this blog post on DDR deployment where I surveyed the deployment of the standard among open resolvers. The conclusion was that we have roughly 300 k open resolvers that return DDR records, But the great majority of them, over 90%, would return DDR records of Google DNS, Cloudflare, and other big public resolvers. So it looks like the great majority of DDR implementations are incorrect, but the open question was why is it the case? And the hypothesis I tested for this paper is that what I was observing was transparent forwarders. So those systems, they did not generate the responses by themselves, but what they did was that they took my DNS request. They forwarded it upstream to, for example, Google DNS, and then they got the response and forwarded it back to my client. So the tricky part here from the measurement point of view is that as a client, as a scanner, we do not really know which entity generates those DNS responses. So what I did this time was to repeat the the very same DDR measurements, but then what I did on top of that was to also perform the the forwarder detection measurement. So that was a completely different scan independent from the from the DDR scan. But what I found is that the great majority of nonverified DDR deployments are due to transparent forwarders. So the RFC does mention these cases and it says that such queries for resolve that ARPA should not be forwarded upstream because the received designations would not be verifiable, but that's what we indeed observed in the wild. So the two takeaways here are that the great majority of what looks like DDR deployment is transparent forwarders. But otherwise, among open resolvers, we have roughly 900 systems that actually deployed DDR correctly. Thank you.
[00:53:14] Georg Kot: Thank Thank you. We got first question in the queue. Marco?
[00:53:18] Marco Davitz: Yes. That's me. Marco Davitz, S. A. D. N. Thank you for this presentation. I have a question. Did you also look at client behavior? Because I've been trying to configure my open resolver with the proper records, and I have never been able to see a browser or whatever switch from ordinary DNS to the encrypted DNS. Apparently, I'm doing something wrong. I have my IP addresses in the TLS certificate, so I'm not sure what I'm doing wrong.
[00:53:47] Yevgenia: I did a survey client deployment. I'm aware of certain operating systems and DNS software that do support DDR, but I'm not aware of well, I did not perform any measurements to see whether somebody is actually doing it in the wild.
[00:54:02] Paul Schmidt: K.
[00:54:05] Tommy Pauly: Alright. Tommy Polly from Apple. Thank you very much for doing the measurements here. It's great to see that. Sorry. I think overall, your findings make sense. I was curious if the measurements you were doing could also help detect cases where we would expect the actual records to come back from the public resolvers. Let's say the user has locally configured themselves to point to quad one, but something along the network path may be dropping or intentionally blocking the discovery of this.
[00:54:42] Yevgenia: No. I don't think I can do it with my measurement setup.
[00:54:56] Unknown (Firefox Developer): Thank you for this research. We were just starting to implement DDR in Firefox at the moment. And I think this is very valuable research in deploying it because knowing whether a network actually supports it and whether it is actually likely that the certificate will match is very valuable. So thank you.
[00:55:16] Yevgenia: Thank you.
[00:55:19] Georg Kot: Then we have you go ahead.
[00:55:21] Lohit Karyala: Thank you. Very interesting. Can you say a little bit more about how you did the measurements and how many repetitions you did in order to
[00:55:30] Yevgenia: So I have been doing these measurements for three years already, and the the results do not change at all. It's basically a straight line. So I do it once per month, once per week. It doesn't matter how often you do it. The results don't change. I'll probably stop it.
[00:55:49] Lohit Karyala: Babur Patch by Hasuplatta Institute. So we have as independent researchers, we have also done similar study and can also confirm similar results what has also found.
[00:55:59] Suresh Krishnan: Thank you. Thanks.
[00:56:01] Georg Kot: Okay. Thank you very much again.
[00:56:03] Suresh Krishnan: Thank you. Thank you. Jayshree, if you like, I think Jayshree might try to present. Hello. It's co author.
[00:56:14] Georg Kot: Yeah. I think it's upper time. Okay.
[00:56:16] Jayshree: Yeah. Hi.
[00:56:17] Marco Davitz: Four minutes. Four minutes. Yeah.
[00:56:19] Georg Kot: It's good. Okay. Let's give it a try.
[00:56:24] Jayshree: Hello. Am I audible?
[00:56:26] Suresh Krishnan: Yeah. You are.
[00:56:28] Jayshree: Yeah. Okay. Okay. So I can start. So this work tries to actually look into the different quick based protocols, specifically DOQ and DOH three from the perspective of from the perspective of how they can perform specifically in Indian networks. So the study is mainly focused on India, and we have tried us tried looking into it from various different networks. Like, the most two popular mobile networks, that is Airtel and Jio, these are the most popular ones available in India. And apart from that, we have also tried to check it via the the stable university Wi Fi network. So as we already know, that Quick offers several benefits like real world deployments and or zero RTDs and so on. So while there has been a a a large span of studies on all of these and the most recent ones that came up in, PAM twenty twenty six, we were mostly interested, trying to look at it from the perspective of the Indian network. So the main question that we tried looking at was how deployable and performant are encrypted DNS protocols and real world heterogeneous networks when India is a very big player, from the mobile and the broadband perspective. So our objectives has been to look into look into the real valid encrypted DNS resolvers of India and by measuring basically the connection setups, the zero RTPs, and the steady state latency across the different networks. So we also try to analyze the the availability of resolvers in India currently. So next slide, please. Yeah. So this is the rough methodology. So we this is something that my student developed, so he has a better understanding of it. So we can take over questions later. And so we basically developed to go based tools for our experimental setup. So one is QDNS scope for discovery and validation of the resolvers using ZMap scan, and we are mainly use the APNIC based dataset as well as the GeoLite address space to develop the to to find out the resolvers that are currently available in India, and then we, we looked into we used, QDNS perk for the performance and benchmarking analysis, and we mainly looked at the cold start latency as well as the zero RTDs. As I mentioned, currently, our study is limited from the fact that we have only tested from our university itself, so it's a single vantage point study. And roughly both of these spans roughly for a week. So both the discovery as well as the performances from that aspect. Next slide. Yeah. So this, figure currently shows the adoption trends in India. So we only found out approximately 28 validated encrypted DNS resolvers where we identified, and this we identified across 13 major ASNs and 18 ISPs. And, out of these, 18 out of 20 resolvers were supporting DOH. And since we only tested, for encrypted DNS, we did not obviously look for, DO UDP or DOTCP or DOT. And out of those, we found a very, very small number of resolvers supporting DOH three and DOQ, and they roughly stand out to just two resolvers out of this. And 82.1% of the resolvers out of all the resolvers that were validated, we found that, well, they had valid certificates, and the remainder were using certain self signed and expired private certificate. So if we look into the figure, you'll see that DOH three is basically two number of resolvers, whereas DOQ is 13 resolvers, and DOH is roughly 20 resolvers. And most of these are towards, towards the Western side. Some of these are towards the Western and South Southern sides of India with very, with almost no deployments along the Eastern and Northern sides of India.
[01:00:49] Suresh Krishnan: Thanks, Jayshree. I think we are out of time. So do you wanna summarize the Yes.
[01:00:53] Jayshree: Just next slide, please. Yeah. So this this is what we currently see that DOH is widely deployed. However, DOQ adoption is moderate, and DOH three remains largely limited, indicating that transition to quick based DNS remains almost incomplete when you look from the Indian perspective, and quick based protocols do show, the clear performance benefits, compared to DOH. But, of course, we did not see it, with DOUDP because this is what largely has already been studied. So, these are some of the limitations, and, we would look into it and, obviously, broaden our study in the coming days. Thanks.
[01:01:33] Suresh Krishnan: Thank you very much.
[01:01:34] Georg Kot: Thank you very much. Any quick questions?
[01:01:47] Suresh Krishnan: I did have one, but it's answered in the future work. So I was thinking about, like, you know, they had only one vantage point for this. I was wondering if, like, you would expand to multiple vantage points, but it's listed in the future work. So thank you very much, Jayshree. Thank you. And thanks, Rohit, too. Yeah.
[01:02:01] Georg Kot: Thanks again.
[01:02:06] Suresh Krishnan: Thank you all.
[01:02:06] Paul Schmidt: I want
[01:02:07] Suresh Krishnan: to to stay between you and lunch.
[01:02:09] Georg Kot: All our speakers. We won't hold you long more for lunch. Be back at 02:00 for measurements.
[01:02:14] Suresh Krishnan: Thanks, Yael, for doing the session. Thank you.
[01:03:12] Marco Davitz: If you if you are nice and and
[01:03:15] Suresh Krishnan: How did they give it to you? How they give it to you?
Session Date/Time: 20 Jul 2026 07:00
[00:00:26] Maria: Thank you. Okay.
[00:01:19] Oliver Hohlfeld: You start. You start.
[00:01:21] Maria: I thank everybody. Please get settled. We are getting started soon. And, Colin, if you can get the door behind you. Thank you. Thank you. I'll get kicked off.
[00:01:57] Oliver Hohlfeld: Okay. Welcome everybody to this year's ANRW workshop. In addition to the series of workshops we have since a couple years. This is a joint workshop of the the IRTF and ACM. And Maria, As it is an IRTF event, the note well applies. So if you have any IPR to report, the IRTF chair is right in the front. Also note that we have audio and video recording. So, if you're presenting or if you go to the mic, you will be recorded. Code of conduct also applies. And please note that if you participate here, go on to Meetecho. So that should be probably the next. And oh, no. Wait a second. Yeah. You should go log in to Meetecho. And if you want to speak, go in into the queue queue and then go to the mic. Or if you're online, we will bring you in.
[00:03:17] Maria: Yeah. And if you're in the room, if you join the Meetecho, use the on-site client or lite client so you don't have to have audio and video. So we use that for queuing. And if you cannot get on, we'll be fine, like, to kind of let you be in the queue, but, like, please make sure that it's easy for us to sequence anybody who's remote and also local. Thank you.
[00:03:37] Oliver Hohlfeld: Yeah. The goals of the IRTF, to repeat, are to support long term research related to network standards and protocols. This workshop is also supportive in this area. As for this workshop, we had 61 submissions, out of which 48 were long papers and 13 were short papers. We had a regular TPC selection process, which led to the acceptance of 20 long papers and six short papers. Overall, 26 accepted papers, which we will experience today during the day. The sessions are mixed. They are organized according to topics. And we will end the sessions with appropriate short papers so that we have little bit of mixed flow here. First of all, I want to thank the TPC. So, here are the TPC members. It was a real pleasure to work with this TPC. It was a rare event that almost all of the TPC members did their reviews in time, actively discussed papers online so that we had a well informed and engaged selection process. I hope this will be reflected in the Pragueramme. Online material. The papers are, at least as of yesterday evening, not yet online. This is due to a delayed process by some complications with authors. So, I asked ACM whether when they will be able to put this online, but we had on the last minute some issues with individual papers. So, that's probably the reason why it will be delayed. We will link them on the website as soon as they are available at ACM. Yeah. Again, sign on to Meetecho and join the queue if you want to speak. There
[00:06:07] Maria: there is one. And if not, there's a QR code in there, like, right next to the mics that you can use to join in. And The
[00:06:15] Oliver Hohlfeld: The ITF Pragueram has a link. Worked. At least I used it. For the agenda, we are in the welcome already. We have a keynote which we will announce in the second. Then our first session will be about network management. We have a break, so we aligned with the schedule of the IETF. Then before lunch, we have the network security and privacy session. After lunch, we will have a longer measurement session and another longer session about transport and media. And this will conclude the workshop. So, as for the keynote, I am very happy to introduce Johanna Ullrich. She is from a newly founded university here in Austria, but she will introduce this university herself because she knows better than I. Johanna Ullrich has worked extensively on Internet protocols and in particular on security and privacy of Internet. So I will welcome Johanna And So that should work automatically, right?
[00:07:33] Maria: Yeah. It's just the slides.
[00:07:34] Oliver Hohlfeld: Yeah. Okay. So we bring your slides up.
[00:07:37] Maria: Yeah. Yeah. No problem.
[00:07:41] Johanna Ullrich: Should I use this mic?
[00:07:42] Oliver Hohlfeld: Yeah. Yeah. Yeah. Yeah.
[00:07:47] Johanna Ullrich: Is it?
[00:07:48] Oliver Hohlfeld: It's a little bit higher. Okay.
[00:07:58] Johanna Ullrich: Yes. Just can Can you hear me? Oh, it's good. Isn't it?
[00:08:03] Lars Eggert: Yeah.
[00:08:04] Johanna Ullrich: Well, okay. Thank you for the for the warm welcome. As already mentioned, my name is Johanna Ullrich.
[00:08:11] Speaker 4: I'm not working.
[00:08:12] Johanna Ullrich: Oh, still not?
[00:08:13] Anudeep: I'm getting close. This one?
[00:08:17] Johanna Ullrich: No. That doesn't work. Is
[00:08:19] Anudeep: it on? It's on.
[00:08:22] Johanna Ullrich: No. Now it's better? No. No. Okay.
[00:08:32] Oliver Hohlfeld: Can
[00:08:35] Maria: you hear me?
[00:08:39] Johanna Ullrich: Okay. Now it's better. Well, okay. Thank you for the feedback, but maybe then we get rid
[00:08:45] Jefferson Campos Nobre: of that.
[00:08:46] Oliver Hohlfeld: Just switch the microphone off. Okay.
[00:08:49] Johanna Ullrich: New start. Second trial. Yeah. Welcome to my session for the introduction. Now we have the next technical issue that could go wrong. What?
[00:09:01] Maria: Think it should work like you.
[00:09:05] Johanna Ullrich: Okay.
[00:09:09] Maria: Let me take back the control.
[00:09:10] Johanna Ullrich: No. But
[00:09:11] Maria: Johanna, I'll I'll move the slides for you.
[00:09:14] Johanna Ullrich: It's now it should be on. Oh, yeah. But
[00:09:16] Jefferson Campos Nobre: now it
[00:09:16] Johanna Ullrich: goes to six.
[00:09:17] Maria: Yeah. Let me try to give back the control to you.
[00:09:22] Johanna Ullrich: Okay. Try now. Fine. Yeah?
[00:09:24] Maria: Try now?
[00:09:24] Johanna Ullrich: Okay. It's also the from my experience, the more engineers are in the room, the longer we need for the for the equipment. Okay. Thank you. So short introduction of myself. I worked a lot for a local research institute that's here in the city, so SBA Research, and then at University of Vienna, and now I switched to IT:U Austria just a year ago. It's not in Vienna. It's in Linz. For those who have never heard of the city, don't worry. It's, like, one one hour driving to the West Of Vienna, and it's basically everything that's not Vienna. It's smaller. It's more industry, but still it's nice. It's nice. What why I introduced the university here because it's kind of different from, I guess, what you most of you are, let's say, used from a university. First, it's a new public university, which is very rare. That's that's the new one is established. It's just three years old. We celebrated last week, and the focus is on interdisciplinarity. So on the one hand, we only have research groups that are centered around interdisciplinary topics. Usually, one is computer science and the other one is something else, so that we have still a language that we can speak about. And the second that's the difference is that all our students learn in in projects and only in projects. And at the moment, we only have master's PhD students, but the idea is that we have groups and of students. And in these groups, there is only one, let's say, sociologist and one mathematician and one computer scientist and so on. And together, they should learn. The idea is to find people who can drive, digitalization and transformation in the upcoming years. But with that, enough about IT:U, then, there's the cliche that all of all the Austrians and, you know, I'm heavily bound to these cities. We're kind of tourism directors, and we advocate our city where you're already nodding your head. If you want to get rid of all the cliche stuff, go to the Central Cemetery in Austria. It's very nice there. First, because there's the graveyard of Ludwig Boltzmann, and second, it kind of comb the city kind of morbid charm with a very large cent cemetery, but you can also ask me offline later if you want to have any further tips. And with that, I'm already coming back to my topics. My topics are somewhere between the Internet security measurements and the power grid. So I basically always circle around networks, and this is also how the talk will be. I will always come back to networks, but it will circle around, and I hope you get something out of that because that's the kind of, let's say, other way of seeing it. We start with the first the first part is basically that I want to focus on transitions on the Internet and the fact that the Internet escaped the lab. I mean, this is nothing new for us, and we always see transitions. And maybe most of you will now think about the the I p v six transition that's more on the technological side, and I could also I also worked on that. But I want to more focus on the, let's say, how we use the Internet and what comes from that. And with that, I want to to go through the first part of my slides, and then the second will be about how do we measure that part of the Internet and what should we focus on when we measure before eventually coming to the power grid and why it's also relevant for for the Internet per se, but also not only in the common sense that we are used to think about it. Okay. So the first transition, and maybe that's the oldest one, we all know that the the Internet, was basically a a thing of the university at the beginning. So they were all, let's say, good people. And if we need to punish someone like a student that was misbehaving, we could somehow directly register their account, and everything was fine. That shifted with with the commercialization and that basically everybody went on the Internet. There was suddenly money involved. And as soon as there is money, there also comes somebody who wants to get this money in some way. And and the problem was on the Internet, I mean, there's the famous sketch that on the Internet nobody knows that you're a dog, but also nobody knows that you're kind of a bad person. And this fostered all the security testing and some standards. And it's also what we see here at the IETF. I remembered when I started here reading my first standards and drafts, there was already the security consideration sections, but in most cases, that what was written there was something like, this RFC doesn't consider any security, so there's nothing. Yeah? And there and we have a lot involved in that way. It's just by counting that words, that are written there. It seems that we we we we focus more on security, and we also have now a process that considers security. So I would say there, it's not finished yet, but as let's say say that we see the issue here and, we work on it. Also, I want to remember it's just word count. You know, you can say a lot with many words and little, and or nothing with many words or just a bit with few words, but just as a first first guess or first indicator. So but I think the second transition is essentially one that's not that obvious for us. Usually, technology, no matter whether it's the Internet or the power grid or the elevator, it starts at an add on. You don't really need it. It's a nice thing to play with, but it's not if it's not working, it's not an issue. That was the Internet in the nineteen nineties. In the nineteen nineties, we still would know, if the Internet fails, what to do. And, for example, if we want to reach somebody, we could go, take the old just a regular phone and call. If we want to know what's going on in the world and, online TV is not working, there was still broadcasting and radio and so on and so forth. Yeah? The older among us was to say, okay. We we we know how to do this, with the younger generation, and I also see it with my son. I explained to him the the concept of radio, which was totally crazy to him. Yeah? That we had to wait and go for the right frequency and and all this stuff. Yeah? And in addition, I mean, young people to use it, that maybe this the the easier thing. The harder challenge is is actually that we experience an IP convergence. So we put everything on the Internet and phase out the old devices. So if we look around, who of us has still a very cable bound old phone at home? Hands up. Well, either you all are all asleep or nobody. So barely of us. I mean, maybe somebody has an old TV, but for me, I wouldn't have a TV at home as, you know, an old TV that where I could regulate broadcasting. So this means the Internet is now critical also in disasters. Because when you think about what happens in in disasters or crisis, the advice, at least that's what I learned back in school, you start the radio or start the TV and look for official commands or come official information. If the Internet is broken and all the other things are not there anymore or I can cannot use them anymore, and it's just out. Yeah? And this is essentially the the the second side
[00:18:05] Jefferson Campos Nobre: of the
[00:18:05] Johanna Ullrich: coin of digital convergence. Obviously, we do it because it makes things easier, it makes things cheaper, but we have the drawback on the case of prices. So and we all seen that in the last years. I mean, there were plenty of accidents where somehow DNS was involved And just because somewhere in the world, DNS had some issues no matter what, a lot of things didn't work anymore. I remember a time when we were sitting in the office. Our Internet was working fine, And with every meeting on the meeting app, you didn't know whether it works or not or whether it has some subdependencies on certain parts of the Internet or not. But we also have it in, in events of crisis, so this is one of our measurements works. Long story short, what you see here is, on the list is the different regions of Ukraine. And on the x axis, you see the dates essentially since the start of the war in Ukraine, and you and you see the Internet outages in the different regions. And, you obviously, this is an event of, let's say, extreme crisis, to put it mildly, but you see that you have outages, in Internet outages due to these events, and this is the reason when you have to come back to all other infrastructures or you have to maintain these infrastructures very, very well. Just one alternative infrastructure, for example, that was used just at the beginning of the. There are still some short wave installations not that far from from here that are leftovers of the Cold War that were used during the beginning of the war to get broadcasting to Ukraine, but the luck was also that they are still there. They are not maintained anymore. So in a few years, they will also be gone. So just to indicate that we just lacked this backup infrastructure, and we will even lacked it or we'll we'll lacked it even more in the future. And another the other coin is this is a post stamp from Ukraine, and, you know, you don't have to be a collector on that side. But what they started is to value the most important professions that they need during the war. And if you look at the upper left stamp, that are the people who are running the Internet or repairing all the cables when they are brought than when they when they are gone. Yeah? So you see, it's so important that these people are honored by in their own stamp. Also, this one on the upper left, upper right, that are the people working the cellular network, which is obviously also part of it. Good. And so and this criticality was not there before, and I think it have to. And the first transition that we currently face is or see that the Internet is no longer borderless. That was also I think maybe I'm the the youngest who still know these times that we just thought the Internet will bring peace and freedom to everybody, and this story unfortunately worked out. It was a nice story. It was somehow together with the end of history and so on, and it's just not there. There are firewalls. There are censorships. We have the power of hyperscaler, and especially for those from the European Union, they are aware of this discussion. We're currently heavily debating about getting rid of hyperscalers that are US based, which are practically all hyperscalers, because we feel that there might be the potential to use this power against us in some way or the other. And, also, there are increasingly countries who try to disconnect their people from the Internet. Like, for example, think about Russia where now mobile networks are disconnected or, also Iran or Myanmar in the past. So with these three transitions, we essentially have a different Internet than what was originally planned, and we also have to think about new new questions on what to do with Internet. Because I guess we can agree in here that we somehow want to maintain because it's comfortable, it makes fun, and by the way, we all would be without a job if it wouldn't be there anymore. So I guess we have a strong motivation to keep it. Here, just one that's one of the measurements where we focused on the fact whether TOR is still usable in Russia because then the censorship starts quite heavily, And it seems like it's still working good, but this isn't just as a kind of depiction of of what I just said. And this is just another fact that you see how how serious the problem of digital sovereignty is within the European Union. This is Austria. And all the red dots that are community communities that are currently dependent on Microsoft infrastructure, And Austria is by far not the worst within the European countries, just to get one impression. And this is now the push that we want to essentially get rid of and find a solution for that. Okay. If but if we feel like this is unique for the Internet, actually, it's not. Yeah. My second big, research topic is the power grid, and we see all three, all three challenges, problems, also there. We see on the one hand the geopolitical tensions. Just to be aware, there were some European countries who just reconnected to a totally different power grid, namely from the former Soviet power grid to European just for being safe and being able to run this critical infrastructure. Also, this power grid somehow feels convergence. We all know that electrification is one of the major major strategies to counter climate change, you know, electric cars, heat pumps, whatsoever. And this makes also the power grid even more important than before. And insofar, the trajectory is somehow comparable to the Internet. And also with the digitalization that's at the moment, there comes there comes the need for security. So it's maybe the inverse history from from the timeline perspective, but you'll see all these three points also there. Apart from the fact that also the the the power grid is decentralized, meanwhile, it's really large, and both infrastructure have some ossification or experience ossification. In the power grid, you see because the the infrastructure is large, it needs a long time to build new one. But, also, the Internet feels very shiny and changing on the upper on the upper levels, but we know down there. And I feel you are more the people who are down there, they know that things change very slowly. Just think about how long it took to reach, let's say, 50% of I p v six adoption on the Internet. Yeah? So so we need to monitor and guide this transition, and here, that's also an aspect where we researchers come in. And we need current insights on on the on the Internet, and this is usually what we do with Internet measurements. The insights that's the main purpose or one of the major purposes of this research is to get information or ideas for standardization. The other thing, and I think this will become more and more important now, is policies and laws and, obviously, security testing and also creating automated security solutions. And the last one is this Internet as a center challenge. If you think back about what I've the picture that I've shown about Ukraine and where is an Internet outage and where not, and most of them are related to power outages. So we could use the Internet as a sensor to kind of see what's going on in the world. But if we do that, we need essentially better measurements. And with that, I'm already coming to my second to my second part, and it's Internet measurements and how we do them. And, I put I criticized my own research work because I think it wouldn't be, very nice to criticize others, but I think it's more, but it's more fundamental to the field, how we measure, because at the moment, I feel we get a lot of results, but as soon as we want to act on that, they don't feel accurate enough to put them forward to policy, for example, which means, which in turn means that we have to go back to our foundations and how we measure and what we measure and how what are the limitations and how accurate are our measurements in order to be able to forward them. In the past, I feel this was not, that much, important because on the one hand, for example, say for standardization, I I know, you know, I do examples. They are oversimplified, but just to get to make the point. If you want to know whether a security feature was out of the Internet or no, a rough number on, let's say, 40% use, it was sufficient because then you say, well, could be better, could be worse depending on what you wanted to achieve, but it was sufficient. If we now go, for example, to law, policy, and so on, they they need more, they need more accurate results. Especially now, we have certain laws that come into action sooner and will be executed soon, but also the state authorities have potential to scan or at least they have the the power to do so whether they will be feasible to do it and capable and desirable to do it, that they scan the Internet, and then they act upon it. If you now think that, at least that what I learned, that the state has to be very, very, let's say, careful in what they are doing when, for example, finding somebody because of lacking security something, The data just needs to be more secure and not just, I believe you have this and that. I put you a fine. This might destroy, the idea that we have, of democracies and the right of law and so on and so forth. So if we want to play a role in that, I feel that we have to become more rigorous first. And I just want to give you a very basic concept that helped at least the people in physics some decades ago to kind of improve their their measurement, execution or how they collect data. But just, be aware that I'm also just starting into the direction, so I don't have the world formula or the solution. I just provide you an insight into that. So, basically, in science, you have the idea of induction. So you infer a general rule from observations. And the other one is deduction, so you have a general rule and draw from that rule on on specific cases. So, I have some examples here. So I wander around. There's a park next to it, and you likely meet swan swans there and say, well, this one is white. This is white. This is white. I guess many of you know the example, then I can infer doing using deduction that all swans are white. The other thing would be deduction that there's a general rule that stated all swans are white, and I I'm I close my eyes. You say there's a swan in we a swan in front of me, and I say, well, it must be white. No no other no no no other choice. There but there was large discussion. Karl Popper, by the way, he's from Vienna, but it's not because he's in there, but it's a coincidence. And he rejected induction, as a scientific method. Yeah? He said just observation, doesn't bring anything. I mean, it it's a good start to serve your environment, but it's not, solely applicable as a science myth a method. Science rather advances through deduction. And with that, you can differentiate actual science. So what we hopefully are doing from pseudo science, like, I don't know, astrology to to name something. And also, because I already recommended the Central Cemetery, his grave is, and this is a very Viennese statement. His grave is Indiana, but at a different cemetery. If you're interested in that, let me know later. But in the end, he said, well, we have to create experiments that that are not so easy that they kind of, you know, confirm what we believe to believe anyway. Because we people are somehow structured in a way that we collect only the the evidence that somehow fit our purpose. I remember my my father once calling me and said, you know, I thought about buying this car, have you seen that basically everybody drives this car? And, you know, it's because he he now has considered buying this type and just ran around the streets on this car everywhere despite it was actually a car type that was at least not that much of a thing here. So we the conclusion from that is that we should do experiments, but experiments that are created in a way that they contradict our or that they are so strong that, if there's contradiction, we have to then we have to remove the hypothesis. Yeah? But if they survive these really strong experiments, it seems like a very good hypothesis, a very good theory to, you know, explain our world, but we always have to be aware that there might be at some point be an experiment that removes all all our insights. Yeah? And a good sign scientist also accepts that their theory might be wrong. So we also have just a temporary knowledge. And if you look to the natural sciences and at this interdisciplinary place where I'm at the moment, yeah, I feel like that physicists are a lot more thinking in that way. They have their hypothesis. They do their experiment. They revise their theory, and they are doing this basically in an analyzed loop to their retirement. Yeah? And with that, they move forward. I feel that we do and yeah. I yeah. Sorry. I'm bit because it's lagging. I feel that what what we are doing and, you know, I I don't want to blame you. I blame myself primarily, and I hope that it somehow makes the the message. It feels like we rather look for what we have as a tool to measure because that somehow limits us. And we start with the tool, then we do the experiments, and then we do the data analysis. And I wouldn't say it's a full stop, but it feels we have very different data points, and we are not really providing them or putting them together. Yeah? And maybe that's not so here I have some some some critics about that. So we rather measure what we can versus what we want. Maybe we should go a step back, and then we have some kind of rule of thumb of Pareto Principle outcomes. These are examples from my own research. So we found that many routing loops. Fine. But what to do about it? Is that is that much? Is it not much? Is there a theory behind what to do it? How can we, you know, base on top of this these results. And the second is also and this is also, you know, becoming a a discussion about our major conference, like, for example, IMC, repeatability and replication. So do we get the results a second time? If not, I'm measuring them, but you measure them. That would be the ideal of science, but it seems like we fail at least sometimes. Yeah? So I would argue that we have that and, I mean, it's basically and, again, I include myself. We have an engineering background, I guess most of us. Yeah? There might be some exceptions in the room, but I would still say we must add engineering functions that way. But if we observe the Internet, it feels rather that we are coming from the methodology point more towards natural sciences despite the fact that the Internet is still, you know, an engineered product, but we're moving more that direction. So we should also go from the methodology and from the thinking more in that direction. And as I said, again, I don't want to blame anybody, but it's because we the one who started the field are the engineers who also created this infrastructure. And and so we still work in this this engineering mindset, which is we try to optimize, validate solutions in comparison to to discover fundamental phenomena. For example, we try to meet specifications rather than than, you know, finding overall theory. We want to fit this machine for that purpose in comparison to university, and so on and so forth. Is the mindset that most of us learn at university, I guess, and I think education is still kind of shaping our minds, and it stays there, and sometimes it sticks there till, yeah, till the end of our lives. Some some notable negative examples are, for example, that we believe that geolocation was essentially a soft topic for a decade, and nobody dared to ask whether the results in there are really accurate. Yeah? Just with this replication, you know, start of replication at IMC, for example, you come to a conclusion, oh, the assumptions are maybe not as good as we thought. And, also, for example, there was along the fairy tale of the scale free Internet so that the connection follow a power line distribution. If you have ever seen a router and how many connections it can have, and I guess there there's a 100% chance in the room, you will see that this doesn't hold the physical background of of the Internet. And, yeah, this as I made already the case, this kind of hinders us as a discipline to go beyond policy, security testing, or doing measurements of the real life via the Internet. So, again, my my motivation is and this is more or less the point where I'm stuck, how to, you know, get these principles into our domain, how to kind of transfer the ideas that physics has to our discipline. As we also have a bit different environment. One drawback is for for example, while physics stays pretty pretty the same, yeah, so the the natural principles stay the same, the Internet doesn't. But on the other hand, we have one advantage and that experiments are just cheaper. If I go to my colleague in the in the neighbor lab, she's spending a lot more money than we and also repeating, experiments is way more challenging for her.
[00:39:22] Maria: Johanna? Yeah. Would you like to take questions now, or
[00:39:24] Anudeep: you wanna wait till the end?
[00:39:26] Johanna Ullrich: Maybe wait till the end. Is that okay? Yeah. That's fine. Thank you. Oh, and with that, I come to the last point. Again, going away again, from how we do measurements, going more into the direction of criticality, and also thinking that only because we have permanently the feeling that we have to protect the Internet and its criticality, but we shouldn't forget that the Internet can also be kind of the threat. I mean, yeah, we're discussing about threats to children and social media, but I'm not now on this level, but also that the Internet kind of has the potential to threaten the power grid. And, so these explanations are by the European, power grid because I'm European, but the American or all the other power grids are except from some terms that might be different and some numbers, nominal values that different physics is the same also on the other side of the oceans. Yeah? Anyway, we have 50 hertz as a nominal value here, and this frequency is not constant. It always shifts a bit around this nominal value, and this is, an indicator whether there's balance in the power grid or not. So if the consumption gets too high, the frequency drops, And, at certain thresholds, we have some emergency routines. You see a list there. You don't have to exercise them in in detail. But if you go from 49 hertz, you see that essentially 10 to 15% of the loads experience a blackout, and at 47 hertz five hertz, all the power plants just shut down for for emergency routines because the turbines would break for mechanical reasons. And what we thought so the idea basically come came over one or two or maybe more beers, like all good research ideas, whether you can kind of destabilize the power grid just by creating a botnet and then increasing the power consumption all of a sudden because then the consumption is too high for for what's anticipated by the power grid operators. And then, hopefully, we can bring the the frequency down to the threshold the thresholds we have just seen before. And this was our model. Just to be aware, this was the model of the whole European power grid, you don't have to be an expert. This is a more coarse trained model, but for the sake of this research, it was sufficient. And what you see here essentially at the upper the upper line is 50 hertz and 49 hertz. That was when the first blackouts come, and you see that you can reach this 49 hertz easily before the time of thirty seconds. Why thirty seconds? Well, thirty seconds, that's the fastest response that the power grid can have. So the fastest reserve so there are always power plants in hot standby, which step in in just such a case, and the fastest ones among them are reacting within thirty seconds. So if you're able to destabilize before the second this thirty seconds, you essentially made it because they just cannot, you know, counter anymore. You're just faster. And this is also obvious because if you think how many moving devices do we have in IT equipment, well, almost almost none. If you think about all these turbines that are in the power grid, there is some mechanical rotation and inertia behind it, so it just takes a while. And interestingly, it also is dependent on situation in the power grid. So what you see here is basically on the left side, it's when there's minimum load in the grid. So summer day, nobody is using electricity because they are anywhere somewhere swimming on the lake. But on the right side, we have winter. Everybody's at home cooking, heating, blah blah blah. And, also, the upper row is kind of classic generation gas turbines and stuff like that. And at the lower at the lower row, you have more renewables. So with more renewables, actually, we become more volatile to these attacks because we reached a threshold earlier. So the one the one in the middle with the frame around it, this was the picture that we've seen on the previous slide. And if you compare, like, if there's more load, the attacker is less successful. Instead, if there's less load, it's more successful. And if they're on more renewables, it's also more more successful. So we have to be aware, and I'm not saying that we shouldn't do more renewables. We, of course, should, but we have to be aware of that. And with that also, the Internet becomes more of a threat for the power grid. And if the power grid down is down, I think this is even more an issue than if the Internet is down. And what we calculated, coming back to the Internet, how many devices you need. Just be aware the European grid spans over the whole continent. So we need disinfections also here on the continent because it doesn't work if the botnet is saying working operating in The US because, that's not important for the European power grid. But back then, we started with, let's say, two point five to ten million infection. And because there are more and more IT IoT devices and, you know, smart heatings and cars and so on, we're now down to one point five million infection somewhere somewhere within Europe, including Turkey. And this number will go down further and further. And we started this re with this research, back, you know, ten years ago. We always heard, like, who should do that? Who should bring our power grid down? Well, again, this this situation just has changed. Yeah? If if we look to Ukraine, the power grid is a target. And, also, if you think about the discussion, especially in European countries, we would never send our soldiers somewhere, especially as an Austrian, we would never do that. But, you know, fussling around with the Internet remotely via cyber operation, that's easier to sell to your public. So this is a nice way to actually do a nation state attack. And by now, we were always just thinking about, you know, is there an adversary behind that wants to bring our network down? But, actually, we don't even need the the part some some adversary that's on purpose wants to bring our network down. We also calculated the power consumption of Bitcoin and Ethereum. I know Ethereum changed the proof of stage, but it's just a showcase that they exceed the reference incident. So their consumption within Europe exceeds the reference incident. So the maximum maximum number that the European power grid can sustain. And, yeah, we can could now say, well, forbid Bitcoin, but this wouldn't solve our issue because it could also be, I don't know, some computing error in or some software error in, let's say, a car that's widely used within Europe. Yeah? And if there are millions of these cars around and they are electrical and charged and they have a software error and they stop suddenly loading in the night, altogether, we have the same issue. We also had back in the time there was, for example, a leap second error in Linux. Thankfully then, you know, data centers were not that much of a big deal than today, and they all increased the power consumption all of a sudden. If that would happen today, I think we would have the same idea. Just here, if you remember Bitcoin, it's just, you know, the example, that you don't need an adversary. A random software error is enough, and we all know that random software errors will be out there now in ten years and also in the future. So what are the the consequences? Well, we have outages, blackouts in one region or more regions. From my experience, if it's just one region, it's typically okay because you can help from the other regions. If the region become larger or all of Europe, Europe, it's really an issue. Destabilization usually takes thirty seconds. Restart two weeks if you're lucky, and then the problem just starts because then you have to slaughter all milk cows because they are maybe ill. You have to to have you have to clean all, cooling facilities, you know, in supermarkets and so on, clean all pipes for water supply. So this is an endless light nightmare. And some studies from the German Bundestag even say that we are then back to postwar economic position. Yeah? So and with that, I always said we don't need an adversary. It's also just the regular Internet as we're at the moment because it could be a threat, and this could also be again a starting point for Internet measurements to see what's out there, what could be reached to form such botnets, or just to cause such such disasters as we have just mentioned. So now with that, I come to the conclusion. While the Internet is a critical infrastructure, if you were not already aware of this before, I hope you're now, the Internet, you know, the Internet is threatened, but also the Internet threatens other infrastructures, most notably the power grid. We need more insights. This is, by the way, not only for the Internet the case, interestingly, also for the power grid because we have the situation that the former monopolies typically have the insight, but all the others who should push the the energy transition, they they are essentially working blind in that field while and we need more rigorous methodology, therefore. And I hope I roughly made time. And with that, I'm happy to take questions. At least I'm already aware that there were some. That's good.
[00:49:38] Maria: Thank you. That That was a great talk. Yeah. Thanks. Go ahead.
[00:50:37] Johanna Ullrich: And I I fully agree with you that this is one of one of also the, let's say, the blockers to kind of changing something. Yeah. Because we have established structures, and we want to, you know, keep them because they are cozy. Thank you. Yeah.
[00:50:54] Maria: Thank you. Curtis?
[00:51:36] Johanna Ullrich: I I think and it might be not not, you know, not a satisfying answer. I think it's the black and white. I think we can learn from from from both. I think what you introduced with the with social aspects, that's for sure true because still the Internet is an engineered engineered structure. Yeah? But still, I think we can learn a lot from the method and the the rigorousness and think twice about measurements from natural sigh sciences. But this was maybe maybe maybe a bit hidden in my talk. It's like, for example, that the Internet changes is really an issue. We could but as as I said, I don't have to I don't have to conclusion. We could, for example, go more when you go more into the natural science direction, that we measure things that are that are constant over time in the Internet. Whatever this is, I have no idea at the moment. Yeah. The discussion is whether can we find something. If we can't make it, obviously, we'll fail. Yeah? I think it's not a black and white, discussion. I think it it needs both. Also, mean, I know that this might not be a satisfying answer, but life is complex.
[00:52:52] Maria: Thank you, Curtis. And, for the people who participate, can you Curtis, can you post the link to the paper onto the chat so that people can have it after the meeting, the hot next paper? Okay. Cool. Thank you. Thanks.
[00:53:07] Johanna Ullrich: Yep. Thank you.
[00:53:08] Maria: Please state your name since you're not
[00:53:10] Anudeep: on the queue.
[00:53:19] Speaker 4: I think everyone agrees that that's really critical. It's critical in the sense in in in the benefits that it brings to us. It's transforming our society in every way. And it's also critical in, I think, in the side effect that it brings to the to the society. Recently, we see more and more cases of government regulations to reduce the minimum age that our younger generations have access to the Internet and maybe to the social network. So in your opinion, can the Internet community play a role in our fighting with, this addiction, the population level, or more in particular, our young generation's addiction to the, Internet and to the social media? Thank you.
[00:54:12] Johanna Ullrich: So so my the scientists in me say as well, that's not my that's people that's mainly people. That's not my field of expertise. But, obviously, we have to like, before our technology, think of how how we want to use it. And and there I don't know whether these bands that, for example, currently applied in Australia, and we're also discussing them, are the right choice, but I don't dare to say yes or no because I'm not the expert in that field. But for sure, we have to discuss the the impact of the Internet on on on society. My point here is more about, you know, the structure, the infrastructure more on the infrastructure part. But as a citizen, let's say, I'm I'm with you. Yeah.
[00:54:58] Speaker 4: Okay. Thank you.
[00:55:03] Maria: Thank you. We still have a couple of minutes left for questions. If anybody has any, please come up.
[00:55:08] Oliver Hohlfeld: So I'll one in between. I mean, thank you very much for your comparison with the natural scientists, in particular the physicists. One of the activities is to design and operate subtle, sometimes large, sometimes expensive measurement infrastructures. And actually, the Internet community has a few measurement infrastructures, just to name Kaido's telescope also, but it's relatively few. There's no professional operation of this. There's no funding for this. And isn't this a race to arms to change our perspective that we have? If we have a critical infrastructure, we need a professional measurement tools and also means to measure. We all know that measurement is always in conflict with the hyperscalers, with the operators, with and so on. So that's a tension thing.
[00:56:15] Johanna Ullrich: Yeah. I I think we need I think the prob or the the reason why we don't have it yet because it's relatively cheap to run measurements. So let's say, not that perfect, but good enough measurements with our with the infrastructures that we have. And in comparison to physics, what we do with our measurements, are not worth worth mentioning the money in comparison to what they do. Yeah? But I agree with you that it would make sense too.
[00:56:41] Oliver Hohlfeld: But this is really good enough. I mean, we have seen in the last ten years or so a lot of Internet measurements where people, sometimes us, sometimes others, actually worked out that the measurements are
[00:56:55] Speaker 4: crap
[00:56:56] Oliver Hohlfeld: because they are so incomplete, the infrastructure is so incomplete, that the claimed observation just does not hold.
[00:57:04] Johanna Ullrich: Yeah. Right. Right. Yeah. Yeah. Yeah. And this is maybe the first thing in all this is getting the measurement method right because then we can if we get that one right and, you know, which observation points do we have, How are the buyers? Knowing the buyers helps us in proceeding forward. I mean, for ex just just a small anecdote. I have a physics student. She just finishes her master's, and she's she wants to pursue a career in traditional physics. This is also why she wants to leave my group. I'm very sad about that. But she was at the beginning she's really nice and and capable person, and she was sitting here and don't you have any error calculation here? And we no. You just put the average you just calculate yeah. Uh-huh. I feel bad when doing this. So yeah.
[00:57:54] Maria: Thanks, Yelena. And I would like to also, like, make a plug for the MapRG at the IRTF. They do a lot of measurement stuff. And if you find further things, keep continually engaged with us and bring the stuff to MapRG, and Doug would greatly appreciate that. So thank you very much for your talk. It's fantastic. And thank you very much for coming, and hope you can stay around for the day.
[00:58:18] Anudeep: Thank you. Thanks.
[00:58:28] Oliver Hohlfeld: Yeah.
[00:58:32] Maria: We have a couple of minutes. I just wanna provide a small intro to the network management session. So we have five papers, and, like, they're, like, extremely varied. So, like, we have quite a bit of topics going from edge clouds to dynamic segment sizing to zero copy load balancing and stuff and so on. So we wanna continually run through the presentations. So we have twelve minutes for the talk and three minutes for questions on the long papers and then four plus two for the short papers. So we would ask you to hold questions till the end of the presentation. And thank you very much for being cooperative and interactive. Thank you. And, Jefferson, if you can come up, I'll get your setup. You want control? Can you see
[00:59:37] Jefferson Campos Nobre: Hello, Willy.
[00:59:39] Maria: Just sec. I'm passing the slide control to you.
[00:59:53] Jefferson Campos Nobre: Good morning. My name is Jefferson Campos Nobri, and I'm going to present the work entitled unified network abstraction layer via knowledge graph, which is a work by by mainly by the master's dissertation from Philippi. So this is the summary of the presentation. We start with some motivation for this for this work, then we have the related work for the concepts employed in this work, then the proposed solution. I'm going to talk more about the information model that we're proposing in this work, some results, and then I will end the presentation with the conclusions. So the heterogeneity problem is a well known problem for a for a long time in considering network management data models. So regarding the ITF, the young is the standard modeling language for network configuration considering the framework with NetConf or ResConf. And one of the issues that is known for a long time is that there is a high structural and semantic fragmentation across vendors. So in this context, couple of years ago, there's this workshop from IAB and then mops, which highlighted that the distinct models that are used by the different vendors hinder network operation automation. Sorry. So talking about this heterogeneity, the question is there the challenge for having these multi vendors using the same model. So for for motivation purpose, we use a minimal case in this in this problem area, which is the admin status attribute. So, for example, considering even even this really minimal configuration, this is represented different across vendors. So in this case, we're using, as an example, Cisco, Huawei, and Juniper. And for the same attribute, they have three different ways to represent it. And in Cisco and Juniper, if you have an empty leaf on the model, means that the interface is administratively down. So in in other in in this example, in Huawei, we have some attributes that is that's not using the empty as any status. So you have one concept, three different implementation, and this heterogeneity is repeated across hundreds of young modules. So it's a very well known problem. So there is some work on on regards of this this problem. So, for example, there are some people that are using a structural conversion for young models to different ways to represent the information. There are other approach that use, for example, some kind of analytics and operation regarding the data models in order to have some form of unified view. And in this work, what we propose is a unified abstraction. So our our approach is try to harmonize different multi vendor implementations in order to have a unified abstraction considering semantics for the models. So we have the contributions of this of this work is has this work has three main contributions. The first one is to have this kind of a a semantic model capable of representing these network concepts despite the difference between the young models and the ways that the vendors implement this kind of of models. The second contribution is to have a three layer knowledge graph. It's some kind of data model information model that is used right now by different working groups and research groups in the ITF. And we want to use this model as a way to preserve the implementations from the vendors regarding the YAG models. And this abstraction can be used in order to have unified semantic queries considering the information model. So in our architecture, we have in the the right side of the of the slide, we have this this image that shows at the top the devices. So in order to have the information from the device, we use a NetConf parser. So after that information is used to build a knowledge graph, and you use for that the Neo four j, which is a knowledge graph database or framework as you wish. Then we use Decipher, which is a query language from Neo four g in order to provide this unified network view. Considering the way that we we try to abstract the information, we organize our information model considering three layers. So the first layer in the top, in the the right side of the the slide, is the semantic layer. And in this layer, we define vendor independent concepts. So it's either the upper part of the abstraction. Then in the middle part, we have the mapping, which maps these vendor dependent concepts regarding specific implementations. And in the bottom of the right side of the figure, we have the operation layer, which stores the actual network values. So we have this organization for the information model because we think that this information can help to have a good abstraction about the about the data models, but also preserves the efficiency of this kind of a query regarding the Neo four j. So talking a little more about the semantic layer. So in this example, we are using the of network interface and interface type. So you can see that it's more like an ontology for you that are aware of this kind of information model. So we have the network interface, and the network interface has several features. So this has in the edge has feature and then other nodes which represent which represents the values that are associated with this this node. For example, the network interface has a description, has operational status, and so on and so forth. So in using these, we can represent some kind of universal concept of a network interface. And universal, as I mean to say, that's not attached to a specific implementation and with different features. So after that, we have the multi vendor mapping layer. So in this layer, we associate the features with several others regarding the implementation. So, for example, for admin status in this context, it's implemented as a net conf query for for this kind of of implementation. So as you can see, we are not using at this time, but it's possible to have different ways to collect the information from from the network devices. It's not only for NetConf, but in this because of this work, we're using just NetConf for that. So we preserve, using this mapping, the abstraction that we, I I I showed in the last slide, but we can use different ways in order to to feed this abstraction. And the operation layer is the the layer that stores the operation starts of this of the network and of the devices in the network. So this reflects what we the the operation and and cooperation behavior which is captured via NetConf. I think that one one thing that I can also say that it in this work, we are just considering configuration. But as we we see that it's possible to use this for different network management tasks in features. So let's see what the the holds for the work, but it's possible to have that. So in this context, I'm showing in the right side of the slide a more, how can I say, holistic view of the the layers considering all the the layers that built the information model? So we implement as I said I said before, we implemented that using Neo four j. So you can see on the on the top of the right side of the slide, the Cypher query, and Cypher is the, as I mentioned before, is the query language, which is used in Neo four j. So we have this query. And using this query, we can have information from three different implementations. And in this context, Cisco, Huawei, and Juniper. And we can provide a unified result considering, in this context, as a proof of concept, the admin status of the interface. So what we can see in this result, and it's a proof of concept, that it's possible to have a unified view of different considering different vendors for some value of of the device in the network. And, of course, after that, we can use this information, for example, to feed models and to to do other stuff regarding network management tasks. Okay. So, for example, another way we we try to understand the the features of the of our proposed solution. So we can use it for resolve young paths from semantic concepts, and this can enable dynamic NetCompRPC generation across multiple vendors. For example, in this context, you are looking for you were looking to resolve a path considering, for example, the the vendor Huawei. So we can have this as a Cypher query in the middle part of the slide, and then the the response, it's in the the bottom type of the slide. So we show that's possible to do that considering our proposal. So as a summary of the the results, we we have a single point of inference considering our work. So the semantic layer provides that defining identified meaning for the for the the the the data model. So we translate the data model for an information model, and we doing that, we add up the semantics for the the data we collect from the devices. The second result we we want to to highlight is the automated discovery that's possible to do, and then show the the graph that use it by by the query in order to solve, for example, young paths and providing the ability for for construct NETCONF RPC payloads. The fourth result is the unit unified state retrieval. So as I mentioned before, we can have a single Cypher query in order to retrieve information from two different three different implementations regarding three vendors in this case. Of course, it's can it can be different implementations for the same vendor, which is something that it's also known in the YANG world. So it's possible to do that regarding different configurations, different models of equipment for a specific vendor also. And the last one is the context of semantics. So we using this semantic layer, we can have, you know, like, raw network values gain, reach a standardized context meaning and other possible properties or features that wants to, you know, to unify with the data models. So the conclusions of the work. So we provide the unified semantic abstraction for heterogeneous E and G models. We proposed three layer information model preserving vendor specific implementations. What's more, we have a way to query these knowledge graphs using as our implementation decipher query language, but this is something that is more related to the implementation of the knowledge graph at this point, but it's something that different knowledge graphs can can be used. We we think that that's possible to have different ways to implement that. But in this proof of concept, we used Neo four j and Cypher query language. And the the final conclusion is that we provided a proof of concept for our work regarding three different large vendors, Cisco, Huawei, and JonyPen, and showing that it's possible to do that. And, of course, we can add different vendors and different models of devices to, you know, to have a larger evaluation in the future.
[01:13:59] Maria: Thank you.
[01:13:59] Jefferson Campos Nobre: So the future just to to to finish the future research directions, want to automate the ontology or the knowledge graph creation using LLMs for automatic parsing of young models and generate from mapping layer. We wanted to integrate with real time network monitoring tools, so we are trying to have, you know, another network management task, which is monitoring in order to to see if it's possible to use our our knowledge graph in efficient way. And we also want to expand the ontology by implementing implementing additional functions for young models targets in this work. So I think this is my last slide. So thank you.
[01:14:45] Maria: Thank you. Yunze. Thank you. Go ahead.
[01:14:54] Anudeep: Hey. Okay. Hello?
[01:14:57] Yunze: Hey. Hello, Jefferson. Hello. I'm Yun De from Chung University. Thank you for your wonderful
[01:15:01] Maria: Can I speak into the mic? We cannot hear you.
[01:15:03] Yunze: Please. Okay.
[01:15:03] Maria: Sorry. Pull out the mic. Louder. Louder.
[01:15:05] Yunze: Oh, louder? Okay. Yeah. And my first question is that since what is the cost of adding a new modular or a new vendor in your mapping layer? And is this process automated or manually? And what is the most difficult part of this? And my second question is that I think that the network computing is very broad. And what is the goal of this mapping to support all the functions or just narrow it to some high value use cases? Thank you.
[01:15:43] Jefferson Campos Nobre: Okay. Thank you. Really nice questions. For the the first question, I think it's it's the the cost. Right? The cost of, yeah, it this is something that, at the first, we are not evaluating because we know that, you know, adding up an information model usually, it's cost on on an implementation. So we're going to do that, and it's parts of Philippi's thesis to have, you know, this evaluation. The second cost, the second one is
[01:16:12] Yunze: The use cases? The broad
[01:16:13] Jefferson Campos Nobre: Use case. Yeah. And, also, it's something that we want to to to push forward in this in the next stage of Philippe's research is to have, you know, a broader evaluation considering because it's really minimal evaluation in order just to prove our concept. But it's necessary to have, you know, different properties, values from the network in order to that maybe not using just, you know, for a device, but try to have the network as a whole. For example, considering Yang and NetConf, and we know that. So we want to have more more properties for the network being, you know, like, implemented and also evaluated and also different vendors because, you know, we we try just a small set, but there is there are lots of vendors that we can use. So the thing for the next steps for Philips' research is to have, you know, a larger evaluation considering, like you you mentioned before, the cost and also, you know, having more complex configurations in possibly other network management tasks.
[01:17:19] Yunze: Okay. Thank you very much.
[01:17:21] Dinesh: Thank you.
[01:17:21] Maria: Thank you. Thanks. Eric?
[01:17:23] Speaker 9: Thank you.
[01:17:23] Maria: State your name.
[01:17:24] Speaker 9: Yeah. And you can hear me now. Good. Thanks. So maybe this is the same question asked a different way. Maybe I still missed it. How much of the ingestion of the model is automated versus a master's student
[01:17:42] Jefferson Campos Nobre: Yeah.
[01:17:43] Speaker 9: Doing work?
[01:17:44] Jefferson Campos Nobre: Yeah. Yeah. I forget I forgot to answer that from from the first. Okay. Yeah. At this point, it's just, you know, like a student work. You know? So for we're just trying to see if the concept is valid. So it's manually we manually ingest this information for the knowledge graph. But we we have some experience right now using LLMs in order to build the knowledge graph. So we think that probably it's it's possible to do that. But, of course, we we have to see it's okay in terms of the, you know, the the the mistakes and the the errors and the how the fidelity of the knowledge graph after using some kind of, you know, like, an automated ingestion. So this is something that it's really no. It is really the the the the point of the the the research, and and I hope that maybe the future, we can present the results considering this part and the others that the other speaker asked about.
[01:18:42] Speaker 9: Good. Thank you. Simi.
[01:18:43] Maria: Thanks, Harrison. Thank you. The line is closed. Thank you.
[01:18:47] Anudeep: Thanks.
[01:18:48] Maria: Done? Okay. Thanks.
[01:18:51] Jefferson Campos Nobre: Thank you.
[01:18:56] Maria: I think you wanna control the slides? Hello?
[01:19:05] Anudeep: Yeah. You can take it. You can take it.
[01:19:18] Tom Guan: So good morning, everyone. I'm Tom Guan from Technical University of Dresden in Germany. And today, I would like to present boosted, a framework that reduce state transfer latency in edge cloud. So the topic to I present today is or got another topic, which is about edge computing. And my framework is proposed to reduce state transfer latency in edge clouds that leverage Praguerammable data planes, like keeping the existing deployments of network functions. So it's worth okay. Let let me start with the motivation behind this work. So network functions are the fundamental building blocks in today's networks, and they are now increasingly deployed in edge clouds. And many real world network functions are in heavily stateful, so they rely on state to operate correctly. And yeah. So just a few example, for example, a firewall that maintains connection information or the traffic monitors that keeps the packet counters. So, yeah, this trend is becoming even stronger with emerging workloads such as AI, where conversation context is another example of state. And in fact, they have attracted a lot of attention from networking research and standardization community for many years. So you can see here. And yeah. So the the previous work also showed us they they consume significant amount of memories and also contribute to many production failures. So this makes this increasingly important to efficiently manage state for network functions. Yep.
[01:21:17] Speaker 9: Okay.
[01:21:20] Tom Guan: So, yeah, So so how do we efficiently manage the full network functions? Let's start with phone tolerance. So you see the the first one here from the left to the right side. The for the phone tolerance, the after a failure, a recovery is the needs the previous state to resume the operation. So the state transfer is also required for scaling. So when new instances are created, so the state is also transferred to them. And the the same applies to the migration when a new instance is moved to another server. So the states moved with it. Yep. So these are just a few example. There are many other operations also require state transfer. And so yeah. So the the the this brings us to the next challenge, the state transfer itself. So
[01:22:21] Michal: so
[01:22:24] Tom Guan: let's take a look at how the state transfer is evolved over time. We started with virtual machine migration. Many people know that, and it's also commercialized regarding virtual machine migration. So for virtual machine virtual machine migration, transferring gigabytes of memory takes seconds. So containers significantly reduce state transfer. Why? Because it's transfer much less state because you see, it it doesn't use the guest OS. That's why we can reduce a significant amount of memory. And more recently, researchers have focused on transferring only the network functions that yeah. Making the state transfer even faster. But this evolution also come with a new challenge. So recently, we have seen many emerging use cases such as smart factories with the robot here and also autonomous driving. These use cases have very stringent latency requirements. So as a result, state transfer needs to be completed with very low latency. Yeah. So, meanwhile, at the same times, the edge clouds also make the problem even more challenging. So for example, here, you can see yeah. The state transfer can also involve in much more complex scenarios. For example, with the split, the network function state can be split across multiple network functions to serve the users more efficiently. Layer, these instances can be merged again to, yeah, when the systems situations change. So this motivates us to design a faster state transfer mechanism. Yep. So based on these observations, we identified three key requirements. First, the solution should minimize the state transfer latency to reduce the service interruption. Second, the solution should preserve conventional deployment of the stateful the stateful network function in cloud environments. Or in other words, existing cloud deployment should continue to work without changing how network functions are deployed and managed. So we make sure the practicality. Yep. And last but not least, the solution should minimize the the overhead to the network functions, making it practical for real deployments. So with this requirement in mind, let me introduce boost state. So we we start from the simple observation. The state transfer between two stateful network function is performed over the network. So this naturally needs to our idea. So instead of relying on the end host only, we leverage Praguerammable data plans to accelerate the state transfer. So let's look at how the conventional approaches here. So the from left to to right. So with the conventional approaches, they perform state transfer entirely on the CPU horse. So this this can preserve the existing deployment, but results in higher latency and and higher overheads. Right? So the next one, more recent approaches moves state transfer to the Praguerammable data planes, so which which can significantly reduce state transfer latency. However, they require deploying network function on Praguerammable date Praguerammable devices such as SmartNICs or Tofino switches. Yep. So this is this make it difficult to preserve existing deployments. Yep. And our ideas boosted here combines the advantages of these two approaches through the hardware, software code design. So we keep the network functions running on CPU host, and we only use Praguerammable data planes to accelerate the state transfer. And at the end, all the three requirements I mentioned in previous slides are satisfied. Yep. Okay. So, yeah, so this is the overview of boosted. So it consists of three simple steps. First, the Pragueram Praguerammable data plane collects and transfer the state to the destination. And when the state transfer is triggered, only the remaining state from the source network function will be transfer via the the conventional software stack. Yeah. That's mean from the source network function to the destination network function. And finally, the the destination network function will reconstruct the complete state by combining state from the Praguerammable network and from the source network function. That completes the state transfer. Yeah. So let me explain the detailed design of boosted. So the the key idea is to partition the network function state into header state and action state. But the header state, it contains five triple information, packet header information. And the action state recalls processing results for example whether the a package should be forwarded or dropped. Yeah? And the the header state is continuously collected and and transfer by the Praguerammable network when the package traverse the networks. And after that, the source network function only transfer the the the the action states to the destination, and destination network function will risk the risk and track the the complete state by combining the header state and the action state. Yep. That's it. And let me so let me introduce our implementation. We implement the Praguerammable data planes using p four on a NetRoom SmartNIC to accelerate the state transfer. And we our stateful network functions run on standard server equipped with Intel Xeon processor and 128 gigabytes of memory. And to preserve the conventional deployment, we develop a portable library to extract the state at the source level function and reconstruct the state at the destination. So we are in addition, we also includes several practical optimizations such as non blocking state reception at the destination to further improve the efficiency. And we evaluate to representative scenarios, splits and merge, and I mentioned earlier. So for split, the state is transferred from one network one network function to multiple network functions. And for merge scenario, multiple level functions transfer their state to one single level functions. And for each scenario, we evaluate two common operation, scaling, and phone tolerance. And we break down the state transfer latency into three components, serialization, transmission, and deserialization. And let's first look at the performance of merge scenario under scaling operations It can be observed. Boosted significantly outperforms the state of the art by leveraging Praguerammable data plans. Notably, it reduced the transmission latency by 60% compared to your previous work. And the same observation, yeah, can be c four phone torrent operation. So boosters also achieve much better performance compared to the state of the art work. So for the conclusion, we presented in this paper boosted that leverage Praguerammable data plans to accelerate the state transfer in edge clouds. And I would like to emphasize that boosters preserves conventional cloud deployment without changing how network functions are at, yeah, deployed and and managed while we can still provide the low latency state transfer. And we implement boosted using p four on our network's SmartNIC. And the results of those boosted achieved much better performance compared to prior prior work. For example, it includes up to 60% lower transmission latency. So by that, yeah, so this is my presentation. So I'm open to questions from audience.
[01:31:18] Maria: Thank you very much. Thank you. Since there's nobody in line, had a question for you. So, one of the cases you handle is failover. So were you doing some continuous state transfer to keep the failover? And if so, did you quantify the overhead for that?
[01:31:38] Tom Guan: Yeah. So make sure that's yeah. Because so in in the paper, we we mentioned two common state operations, scaling and phone tolerance. But I have to to to emphasize that there are many operations such as upgrading the the the software or the network function. This is the common behavior in software engineering and also in in cloud for the cloud infrastructure. But come back to the corrections. So for the phone tolerance, the challenge is us. Regarding stateful network function, we have to to to copy the state periodically. Yeah. Because we don't know when the failure can happen. Right? And this is the challenge. And that's why, yeah, low latency state transfer is very important. Just imaging does in in the commercial cloud products. Right? There are many many stateful network function. And if we have to provide the the failover mechanism, so we need to make sure state transfer is efficient. Yeah. And that's why we we we use Praguerammable data plans. Yeah. Because nowadays, there are already many products, not only in SmartNICs, Dolby no switches, even for Cisco. We also have in our test path so that we can run containers on top, yeah, for the network computing. That's why we can leverage instead of rely only on the the CPU host. That's introduce a higher latency, as I mentioned, and also higher overhead. Yeah.
[01:33:02] Maria: Perfect. Thank you. Thanks. Thanks. Go ahead. State your name because you're not on the line.
[01:33:07] Michal: Hello. Thanks for your presentation. I'm Michal from. So my question was, first, I wanted to ask maybe more about the motivation. So you said that there's more complexity in edge clouds. Mhmm. I wanted to know how it compares to normal clouds. Why is it more complex to do it in edge clouds? What's the what's the reason? And maybe second question more technically. So what I gather from this is that the the NF state, you basically split it to two parts, and part of it is kept on the SmartNIC, and this is what makes it faster. Is it a good characterization, or have I missed something? I wanted to just, like, summarize what I understood from this. Thank you.
[01:33:45] Tom Guan: So first yeah. So if I understand correctly, the question is about the complexity of edge clouds. Right? Yeah. Yeah. So so the the the question is the thing is why it makes sense? Because now nowadays, for example, many many applications such as ARVR, right, or we talk about the Tactile Internet, these applications require very low latency. And often, there are many research where people want to upload the task to the remote server because the server have a more compete committing resources. Right? And that and and for the remote communication, the latency is important. Right? And then yeah. So people want to move cloud resources closer to to the devices such as the robotics. Yeah? But but in this sense, the ash infrastructure become highly distributed at various locations. Right? It can be in shopping center, or it can be at the next to the base stations, for example. Right? And and that's why it, yeah, it makes sense to to make the study in in Edge cloud. And the second questions are sorry. I I forget the second question. What what is it about?
[01:35:00] Maria: I think you should take it offline with him because I think we are done done with time. Thank you. Thank you very much. Thank Dinesh. If we can step up.
[01:35:06] Tom Guan: Yep. Thank you, everyone.
[01:35:17] Maria: Yeah. I give you control of the slides.
[01:35:20] Dinesh: Okay. Good morning, everyone. I'm Dinesh, a PhD student at MPI Informatics, and I'm presenting on behalf of my colleagues from MPI Informatics, Nokia Bell Labs, and from Amsterdam. I'm gonna start with a very simple scenario where we have a TCP connection between a sender and receiver, sender at the left, receiver at the right, and for simplicity, I've removed the ACK packets in here. The connection is working, but there might be some cases that we have a congestion on the network. We have a bottleneck here in the middle. That bottleneck could be a little bit smarter that we have a queue, and that queue has some kind of threshold. And if the packets are passing that threshold, they can be marked with some kind of flag. That flag would go to the receiver side, and the receiver can forward it back to the sender. And sender can decide about the congestion and react to it. But I can change this scenario a little bit to add a little bit more information to the ACK packets, and that is how many congestion flags it has seen since the last ACK that it it was receiving. For example, in the top left ACK packet that you see here, it's just saying that since the last ACK I sent to you, I just saw one congestion flag. For the next one, three, six, seven, and nine. This is some high level idea of the Accurate ECN. It's one of
[01:36:55] Yunze: the
[01:36:55] Dinesh: latest RFCs in the community that maybe some of the authors are sitting here. But the thing is, what we can do with this extra information that we have? We already know that for for classic congestion control, we were relying on the on the packet loss to react to the congestion, but that was what just one signal, packet loss and react to it. Now we have more information. We know to what extent the congestion happened, and the reaction can be adjusted according to that. For example, if none of the packets are having any signals, it means that there's no congestion. If all of them are having signals, it means that we have a very severe case of congestion. Right? Okay. There are some protocols that are working with this accurate ECN, for example, data center TCP. But it is designed for the data center environment. So it's not TCP friendly, and the traffic cannot go to the internet. And also, Prague is some kind of de facto, congestion control for the L4S is also doing this using the Accurate ECN. The rest of the talk is going to be about the Prague. That was the case for Accurate ECN, and Prague has some more features. And just one of the features that it has is that minimum packet window is, limited to two packets, two segments. This is for preserving a minimal packet pipeline and also remaining compatible with the delayed act that that might happen to the DCP connections. But it's giving one more feature to us. That is with more packet, we can increase the resolution of a congestion signal. If we have just one packet in the pipeline, it's either 100% congested or 0%. But if we have at least two of them, it means that it can be 50%, 100%. K? But there might be some cases that we have very severe congestion, very low bandwidth. I'm showing the same experiment, but this time, we have just 100 kilobits of bandwidth here. And the packet sizes are a 500 kilobytes. This means that even for sending a single packet, we will have we will need one hundred twenty milliseconds of time. Right? And the virtual RTT in the in the Prague is twenty five milliseconds. It doesn't matter what is that, but what matters is that in one virtual RTT, we are sending less than one packet. With less than one packet, it means that we have less granularity for the condition, and the scalability of the the Prague won't work at the very low bandwidth. Prague has some solution for this, that is dynamic segment sizing. It's not it's not named dynamic segment sizing, but for this presentation, we will name it like that. And the idea is instead of having one packet that takes one hundred twenty milliseconds, how about we have 10 packets that each one is taking twelve milliseconds? It means that we send smaller packets, but with the same rate. What Prague is doing is for having each segment size, it's using pacing rate and the visual RTT to to decide about the segment size that it's it's, gonna send. But one problem will appear with this, and that is the virtual RTT is noisy. Virtual RTT is some kind of smooth RTT that TCP is estimating, and that could be noisy in TCP, meaning that the segment size estimation could be noisy as well. We want to change this a little bit by using the some kind of dynamic segment sizing that is just relying on the pacing rate, we remove the the the smooth RTT from the formula. So for packets between 100 kilobits per second to one megabits per second, we use segment size between 150 bytes and the full MSS. We wanted to experiment this and see the effects and and do some more experiments. We have one baseline that is all the packets are 500 bytes. Here, I'm showing the segment size in the y axis and the pacing rate in the x axis, and the basings are between 100 kilobit per second to one megabits as I told you in the previous slide. We can have linear segment sizing as well and we can deviate from that, as we already know. For example, cubic is working good in the Internet, so not using linear sometimes help. Okay, so we implemented this in Prog in Linux kernel, we tested this in our testbed. For testing, we are using several experiments. For example, single flow, which is very, very rare in the Internet, but we wanted to have some kind of baseline that is, in a link of very congested link of 100 to one megabits per second, what will happen to the single flows. We had experiments where we have multiple homogeneous flows, meaning a lot of Prague flows are competing with each other. We have cases of competing with inelastic flows. We know that there are small flows in the Internet that don't react to the congestion. So we wanted to show that if if, Prague is competing with them, what will happen to the flows. And last but not least is multiple heterogeneous flows. This is the case that that is happening every day in the Internet, who are competing with other cubic or BBR flows. But because of the time limits, I I don't have time to show all the results. I will go just to the single flow and the multiple heterogeneous flow. The setup is as simple as we had before, sender and the receiver. Both of them are Prague. And now we have some delay emulators on the network that that are emulating the delay that we want to as have as a base delay of the of the experiments. And we have a bottleneck in the middle, which is using dual pi square, the the queuing discipline of the Prague or L4S. And we have iperf in between, and the link is throttled to very low bandwidth, as I said, between 101 megabits per second, another kilobit per second. So what are the packet lenses? If we consider this plot, in the y axis, you see the 100 to one megabits per second of the bottleneck. And in the x axis, you see the base RTT that we are using by the emulator. And this is the default implementation. And it's showing, and the shade of the plot is showing the main packet lens that we have in in our experiment. And it's showing that almost all the packets are around the full one mss. If we use the logarithmic approach, I'm not showing all the approaches that we we implemented, just the logarithmic we see that for 100 kilobits per second, for example, we have very small packets and when we reach to the one megabits per second, the peak one, it's having bigger packets. One of the benefits we saw in this is that retransmission rate decreased a lot. I'm showing the same plot, but this time the shade is showing average retransmission. The left one is the default implementation, the right one is the low guided peak one. And we see that around two times the retransmission of the 20 times the retransmission of the logarithmic one is decreasing. For multiple heterogeneous flows, we repeated the same experiments. This point this time, we have several frog flows are competing with several cubic or BBR flows. And this plot is showing the James fairness index in the y axis, which is between one divided by the number of flows and to one. Closer to one, the better. It's showing the fairness between the flows. And we have experiments of five Prague versus five cubics and also 10 Prague versus 10 cubics. For the case, the second one, each flow almost has just 500 kilobits per second. This is showing all the experiments that we have: fixed, default, logarithmic, linear and exponential. And we see that for the fixed one, the fairness is the minimum compared to the others. And when we use any kind of dynamic segment sizing, the fairness is increasing. The same for the VBR as well. It's very surprising to see that when five frogs are competing with five VBR V3s, the fairness is almost one. But for the case when we have dynamic segment sizing and the bandwidth is lower, we see that the fairness is increasing. So dynamic seg segment sizing is some kind of helping the TCP friendliness of the product. So to summarize, we have seen that we can have different dynamic segment sizings in Prague. We showed that for single flow, it's helping to choose better segment sizes, during the experiments. And we showed that, for multiple heterogeneous flows, not only it's not decreasing the fairness of the flows, but also it's helping to have more fairness during the experiments. You can read more details about the other experiments that we had in the paper by scanning this QR and also I'm available in the job market and if you want to see my CV, you can use this second. Okay. Thanks.
[01:46:39] Maria: Thank you very much, Dinesh, for going through this. And for people who are, like watched the presentation, there's a lot of backup slides too with, like, more information. So, like, please go take a look. There's about 40 backup slides. So thank you very much. And Lars, go ahead.
[01:46:53] Lars Eggert: Hi, Lars. I got Mozilla. Thanks for the talk. Was very interesting. Two sort of, I guess, clarification questions. One is your your flows are bulk flows. Right? Or what what
[01:47:03] Dinesh: flows They are are IPERF flows. They are? IPERF. IPERF. It's bulk. IPERF.
[01:47:07] Lars Eggert: So bulk. Yeah. Okay. Have you looked into non bulk traffic or application limited traffic?
[01:47:13] Dinesh: For the experiment with the with the inelastic flows, were using Hping to to DDoS the the link. That one is not using bulk traffic. But for the product traffic, we were always using the bulk because we wanted to see the reaction and having the flow long enough to have to see the dynamic segment sizing is kicking in or not.
[01:47:36] Lars Eggert: So what I'm asking is I I work for a browser vendor, we many of our flows are not
[01:47:40] Dinesh: It's a very good question. Yeah. It could be one of the experiments we can we can run.
[01:47:44] Lars Eggert: The other question I had was so this idea of segment slicing, it's really old, and it's nice to see it reappear now. There was a Pilc document from a really long time ago, but it talked about exactly this, that you want to have smaller sized packets to increase your packet rate. Have you looked into this? Because, obviously, if send smaller packets, your header to payload ratio gets worse. And I sort of wonder if if it matters here and you've decided it's it's okay or whether that's something you still wanna look at.
[01:48:16] Dinesh: Yes. It is adding overhead. That's not in we cannot hide that. But the the thing is, at least for the case that we had a very congested link, because the retransmissions were were much higher for the for the other experiments, we saw that that the the overhead is some kind of is is a trade off. It's it's covering that retransmission. So we
[01:48:39] Lars Eggert: Because you don't need to retransmit, you can you can justify spending a little bit more on headers?
[01:48:45] Dinesh: Yes, yes, yes. And also, it's helping the scalability of the product in the cases that for a very short term, for example, we have a burst of other traffic and the product is adjusting its segment sizing, it will help the frog to keep its scalability and get enough feedbacks from the network. Thank you. Thank you.
[01:49:10] Maria: Thanks Lars. Gauri?
[01:49:12] Gorry Fairhurst: Gauri Ferhus. I was I I I like this piece of work. It was really good for you to bring it here. I have a question about what happens when we have tunnel encapsulations on the bottleneck. So we have tunnel headers as well as the headers. When you scale down to a 150 bytes, there's going to be an awful lot of tunnel header. Do you somehow figure this out or do you just ignore that? I don't know what to do.
[01:49:34] Dinesh: I think you mean the overhead of the header when we have a smaller packet.
[01:49:38] Maria: I think he means tunneling. So if you have, like, another additional layer of tunneling, the overhead increase becomes more visible at that score's point.
[01:49:45] Dinesh: I think it won't impact this because it's just adding a few a new header. So our packets are still small enough to get enough feedback, right? If we have 10 packets, it's fine. If we have nine packets, it's still more granular than just having one or two packets.
[01:50:02] Gorry Fairhurst: Okay. I'll follow-up later perhaps. That was fine.
[01:50:05] Maria: Thank you.
[01:50:08] Speaker 13: Rohan, this is Keti. How long did you let the experiments run for the fairness?
[01:50:16] Dinesh: I don't recall that, but I think it's more than four or five minutes. So it was long enough to to see the
[01:50:24] Tom Guan: Yeah.
[01:50:24] Speaker 13: Sometimes VBR takes some time to converge.
[01:50:27] Dinesh: Yes. That's That's correct. That's correct. Thank you. We have seen that in previous research.
[01:50:32] Speaker 13: Yeah. But that that should be sufficient, basically.
[01:50:35] Maria: Thank you. Thanks, Roland. Thank you. Thanks, Dinesh. Thank you.
[01:50:52] Michal: Take this.
[01:51:07] Oliver Hohlfeld: K.
[01:51:08] Michal: So hello. My name is Michal. I'm a part of PhD Pragueram at TU Eindhoven. And as a part of my internship, I was doing work on moving layer seven load balancing towards zero copy to get the performance benefits. So first, a small division so we know what they are talking about. So layer four load balancers is where the clients connect to a load balancer, and connections are assigned to the servers, and they are basically pinned. Whereas with layer seven load balancers, what happens is that you as the clients connect, you actually can spread the requests around the servers, and that allows you to do things like content routing. It allows you to have different protocols between front end and the back end, like h t p two on the front end, h t p one on the back end. And both of those nowadays are encrypted. And the problem that this causes is that you have a lot of copies to do, and they are not copies of a easy kind because a lot of headers of TLS, of HTTP are intertwined. So the normal way that this is done is between libraries, there are what we call semantic gaps. The semantic gaps makes it so that when you when one layer wants to ask for the data, it has to provide the data in a in a stream, which makes it hard to to make it into a zero copy solution. But what we identify nowadays is more and more things are offloaded to the hardware. One of those things is a TLS offload, which makes it much more practical to try to achieve zero copy. Also, user space stocks have advanced throughout the years, so we can also try to use them. And you what you actually can do is you can expose the the data buffers to the load balancer so it only uses pointers. It can still translate the headers, but doing it allows it allows the data to stay where it is because it's already received as encrypted. We implement this solution as a as a proof of concept in a, like, a small nghttpx load balancer. It's based on nghttp2 low library. And we use XLO, which is a user space TCP stack. It's very easy to use. It has a lot of features already built in. And to those, we implement the the additional API that allows you to actually receive, as I showed before, pointers, not actually copy the data. And as an experiment, we we run this load balancer. We use just one core, but it easily scales. It's driven to full utilization, so so we are testing how much cycles you are actually using effectively. We just use 10 clients. The requests are load balanced across two back ends, ten ten outstanding request in HTTP. All experiments use XLO, so to make the comparison fair. And we can see that the TLS offload itself only slightly improves the performance. So the cost is not really in TLS. It's in the copies. But when you actually move to to the zero copy, we get at the higher file sizes even up to twice the twice the performance. And as a conclusion, this is a research prototype to show how important the zero copy could be, but the real load balancers, which there are many, would have to move themselves towards getting those benefits. And that's it. Thank you for
[01:55:08] Oliver Hohlfeld: your attention.
[01:55:12] Maria: Thank you, Michal. Any questions for Michal? If not, thank you. Surely. And, yeah, thank you very much, and people can talk to you offline. Thank you very much.
[01:55:29] Anudeep: Thanks.
[01:55:37] Maria: For the last presentation of the session, Anudeep? We cannot point. Nope. No. The remote people cannot see it.
[01:56:03] Gorry Fairhurst: Ah. If you point Yeah.
[01:56:04] Dinesh: Right. That's
[01:56:04] Anudeep: right. Hello? Hello? Hello, everyone. I am Anudeep. I'm going to present about Pinocchio, a framework for Get get the mic a bit closer to you. Hello? This is fine. Yeah. I'm going to present about Pinocchio framework for queuing onset deep prediction in mobile edge networks using reservoir computing. So real time interactive applications such as XR benefit from early queue detection. So Pinocchio predicts this near future onset of queuing. In typical condition signals such as loss or RTD, they have their own set of drawbacks. What we do is we we use queue buildup indicator, QBI, a metric, which is like the difference between estimated packet count that is sent by the server and the number of packets received at the UE. And if this parameter is greater than zero, that means that the queue is building up at the at the base station. And then if it is less than zero, the buildup queue is being drained. So this is our test bed. We use OpenAir interface here for the five g, and we have the edge server which time stamps the packets and send through this OAI core and base station to the UE. And the Pinocchio framework, like, for the detect for the prediction, it is in the edge server. And the UE reports periodically some measure measurements such as QBI here, and we use a certain input features, QBI, and some direct features such as mean variance through a certain window interval, and then train ESN echo state network, and then use the output for the prediction framework. We we train using single flow and then test using unseen TCP cubic and BBR flows. So in this figure, shows the related time from onset. Like, a t equals to zero shows the time at which the queuing onset begins. We want to predict it before t equals to zero. That means that t less than zero, we would like to predict it. And the the green ones here shows some information like that. So there is correlation between one way delay increase increases, and one of the features of this QBI. And we show that, okay, we have some relation which this EcoState network exploits to predict the onset of queuing. And if you move further away from reach further away region, for example, the red ones here, there is no much information here for the ESN to exploit. So at this stage, there is the prediction is not easy to or it's not possible. This is the prediction behavior that we observed. So the actual predict actual onset, like, the event of queuing onset is that shown in circles, green circles, and then the predicted what what we observed through this Pinocchio framework is in red. So we are able to see the we predict the onset of the queuing. We also have the f one scores. So if you want to predict farther from, like, fifty milliseconds away from the event, it's it's okay. It's we can do that. And if you want to do at hundred milliseconds, like, two horizon steps away from this, it's possible. But if it is three, it slowly degrades. These are the conclusions. Next, we'd like to integrate this with the rate controller to see the network utilization and latency performance. Thank you.
[02:00:06] Maria: Thank you. We have over thirty seconds for questions, if anybody has any. Hi. This is Ludwig Deutsche Telekom. Can you please state your name? We cannot hear you.
[02:00:20] Ludwig: Yes. Hello. My name is Ludwig Gagreb, Deutsche Telekom. This seems to replace l four s, or did you investigate l four s? Because I think that pretty much solved the problem. Thank
[02:00:37] Yunze: you.
[02:00:37] Anudeep: Sorry. I couldn't hear.
[02:00:39] Oliver Hohlfeld: You didn't hear. I didn't hear either. Can you speak up a bit?
[02:00:46] Ludwig: Yes. Your solution seems to copy the behavior of l four s, which would explicitly mark the onset onset of congestion and schedulers. And that is deployed already in mobile networks, at least for five g s a. So did you investigate that? Or
[02:01:11] Anudeep: Oh, so if I understand correctly, you were asking about l four s implementation or checking our framework with the l four s implementation. Right?
[02:01:20] Maria: Yeah. Yes. He wants to understand the difference.
[02:01:22] Anudeep: Yeah. So, yeah, we are also working on that. So this is the preliminary results that we obtained. Next, we would like to show that the improvements or any improvements that we get from l four s or other kinds of implementations. It's in basically in the pipeline.
[02:01:39] Maria: Thank you. Thank you very much. And, Nir, you should catch up with them offline in the break. So we have a half an hour break. So thank you very much for being awesome group. Thank you.
[02:01:53] Oliver Hohlfeld: Okay. This is time for a break now?
[02:01:56] Speaker 4: I would like
[02:01:56] Anudeep: to thank that yesterday night
[02:01:58] Oliver Hohlfeld: We'll join in half an hour.
[02:01:59] Anudeep: It's it's one or something.
[02:02:01] Maria: No. It's fine. I was just thinking again. I just wanted to make sure I had everything uploaded in my phone now. Yeah. I'll be back in, like, twenty minutes or so.
[02:02:36] Johanna Ullrich: Hello?
Session Date/Time: 20 Jul 2026 12:00
[00:00:26] Session Chair: Paul Schmidt is remote.
[00:00:28] Benedict: Yeah. Paul Schmidt.
[00:00:30] Danish: That's the same we had before. Yeah. Yeah. That's right.
[00:00:35] Session Chair: Paul, can you hear me?
[00:00:37] Paul Schmidt: Yes. I can hear you.
[00:00:39] Session Chair: Paul, do you want slide control? Like, ask me for slide control.
[00:00:42] Paul Schmidt: Yeah. I have the request up.
[00:01:06] Session Chair: Okay. So welcome back, everybody, to the Applied Networking Research Workshop. I hope you had a pleasant and enjoyable lunch. And you are now fresh and not tired because we have a really exciting session about network measurements, seven long papers and one short paper. And the first presentation is given by Paul Schmidt about PcapML, making network traffic data sets reproducible.
[00:01:31] Anthony: Please go ahead.
[00:01:32] Paul Schmidt: Okay. Thank you. So, again, I'm Paul Schmidt. I'm at Cal Poly, and this is work stemming back many years from when I was at Princeton with Jordan and Nick and who Nick is now at UChicago. And, yeah, we're looking at reproducibility. So likely, many of you in this room have have worked on machine learning and network networking traffic, and you have probably run into all sorts of problems. And this paper, unfortunately, we didn't think needed to exist, but we just keep running into the same problems that we decided to put
[00:02:10] Session Chair: it down.
[00:02:11] Paul Schmidt: So any network traffic machine learning problem depends on labeled traces. So we use this for traffic classification or denial source detection, all all sorts of things. But, unfortunately, dataset pipelines are very often underspecified. And so you've probably had this experience. You see somebody's new paper that has really promising results, and you wanna try it out. And maybe they were kind enough to actually release the dataset. And so you download all of this, and you try to rebuild essentially their pipeline, and you don't see the same results. And the the issue is that the same raw traffic, even when people are kind enough to share their datasets, can end up becoming very different type of datasets, and that's what we're looking to attack here. And so from our point of view, one of the the major weak points we see are that labels tend to live outside of traces. So if people release their datasets, they've got a set of pcaps. And then alongside of those PCAPs are some kind of labeling or metadata datasets, so text files or what have you. And if I see this interesting paper and I wanna reconstruct it, well, I need to join these two pieces of data using sometimes time stamps or time windows, tuples, file names, speech files, etcetera. And the challenge here is that that really small choices become quite consequential. And so what is your definition of a flow versus my definition of a flow? Some people combine unidirectional flows and make bidirectional flows in their datasets. Some have choose different timeouts. We've got different class mappings, and, of course, there's training and testing splits that are not always available for us to see. And so the end result is that papers can compare accuracy numbers against prior work
[00:04:15] Benedict: or prior
[00:04:15] Paul Schmidt: techniques, and they will be essentially evaluating fundamentally different tasks or comparing apples to oranges. And this sort of annoyed us to the point that we decided to try to write this paper. And so in the paper, we go through a handful of these examples of very popular sort of canonical datasets that a lot of people rely on, and one of those is VPN, non VPN. So this dataset's released as a set of PCAPs with the application metadata encoded as file names. So things like email dot PCAP or browser dot PCAP. And that original dataset had 14 classes, so it's seven applications with and without VPN. But the issue here is that that the dataset itself, of course, doesn't enforce mapping for later use, and so papers that came along after the fact derive completely incompatible class sets from that exact same raw traffic. And so in the paper, we do a small survey, and we see class counts ranging from two to 31. And this is not unique to a single dataset, of course. So I mentioned VPN, non VPN, and the root cause there is that these filings are mapped to a single class set. CIC IDS, lots of people use that dataset, but subsequent work found that something like 20% of the labels or or the flows are are simply labeled wrong, and work after that fact has not actually corrected it. And so you end up with a situation where you've got multiple papers working on the same dataset, but they've got completely different ideas of what the dataset actually contains. UNSW is a slightly different one that a lot of people have chosen to model that as a binary class classifier or build binary models versus others have done multiclass using that same dataset, and that's going to change your accuracy numbers by double digit percentage every time. And then, of course, there's DARPA, which has been around forever, and it's known to essentially, the metadata files have all kinds of errors in them, and trying to reconstruct these datasets is is virtually impossible. And so that motivated PCAP ML, which from an engineering point of view is very, very straightforward. There's not a lot of deep novelty here, but we wanted to try to lower the barrier here, and we wanted to build a framework for creating self contained labeled datasets. And the basic idea is we're storing arbitrary labels directly in the PCAP files themselves, and we built it to support two different paths. So the first is with live capture using eBPF, and then we also have offline conversion. So if you've built a pipeline and you just you want to release it as a an artifact, you can do so. And we really just want to stop spending so much time on dataset specific reconstruction. We wanna take these mysteries out of people's hands as much as we can. And so, ultimately, we abuse the pCAP NG comment field, which is available. It's unstructured text, essentially. And so the nice thing about using pCAP NG is existing tools get to still read the trace. So this is a snapshot or a screenshot of of Wireshark with a pCAP ML trace, and you can see the comment field has a sample ID, a process ID, a direction, and a destination. So we just throw these keys in and the data we can grab for any trace, and and we can just put it directly into the PCAP for later use. By default, we tag every packet with a flow level of a five tuple sample ID, and the idea there is that at the very least, we should be able to have trained test splits reporting their sample IDs that were used rather than forcing people to reconstruct using file names later later analysis. So like I said, there's these two paths to creating this trace. The first is live capture. So we chose to use eBPF. The reason being is, well, it's relatively fast, and you can get socket to process mappings directly in the kernel. And so at collection time, we're able to say, oh, this this packet can be attributed to Chrome or this packet can be attributed to SSH or or what have you. We also built a module to parse destination labels using DNS or SNI. And the reason we do that is since we built it as eBPF, we imagine this isn't just going to be running on end hosts. It can be run on middle boxes. So I'll talk about a a small analysis we did. We did ran that on OpenWRT, and we were capturing both at a end host and OpenWRT. That gateway is obviously not gonna have process context, but it still has the DNS and SNI destination stuff. And then we we also built the offline conversion. So, of course, we're not crafting ground truth off nothing. It has to have been captured in the first place, but you can record your construction for later use. And so, for example, you could you could choose how you're going to represent VPN on VPN and then publish that as an artifact with your your paper. And so we did a very small evaluation on this. We did control collection using Linux. We ran five sequential generators. So
[00:09:53] Innocent: first
[00:09:53] Paul Schmidt: was Chrome browsing and Firefox browsing. Those are basically just top of 100 sites that we're going to. We then did Chrome streaming, so we went to YouTube to a number of videos, and then we did SSH and RSYNC to show essentially a failure mode. And we were capturing both, like I said, on the Linux desktop and on an OpenWRT box that was essentially unmodified about running our eBPF mod model. And so the the evaluation is pretty simple. What we're trying to illustrate is what we can piece apart using these different vantage points and and the amounts of the different piece of information we have at each location. So if you have the process label, you can't really differentiate between Chrome browsing and Chrome streaming. But if you have the destination, you can. It's it's kind of obvious. You see YouTube destination domains versus all of the other domains. Conversely, if you're doing Chrome browsing versus Firefox browsing, your the destination labels are not gonna be helpful. You're going to those same sites, but the process label now comes into play, and you can you can peel those apart very quickly. And then the sort of null result or the negative result we had was RSYNC versus Ultimately, RSYNC uses SSH under the hood, and so the the tool we're using to attribute socket to process mapping is just going to show SSH for everyone. So that is a failing. So, obviously, this is not a perfect solution, but we think it's useful. And so one thing we were interested in is a lot of datasets will have these time windows or time stamp based labels, and we were wondering how often that can be a problem. And so we were looking at this, and it turns out during that Firefox trace, something like 5% of the packets were not actually necessarily generated by our doing. Some of those are the DNS Firefox threads, and others were something else. Some of them were from our web driver. Some of them were just random other traffic. And so when you do this, you can you can start to attribute things to the appropriate process, and we hope that this can sort of clean things up and make it easier to train models that aren't necessarily looking at noise. And so just to wrap up this very simple approach, we don't think the packet traces are complete benchmarks if you're forcing people to reconstruct using labels after the fact. So very simply, we store those labels, and you can store arbitrary text inside of the directly in PCAPs at capture time or offline. We don't fix everything. We're not standardizing your task definition. We can't infer things that weren't captured or were encrypted. We've built a handful of labelers for common cases, but the framework's extensible. It's open source and available. And the whole idea is we wanna save everyone time, and we wanna save our own time when we're we're dealing with these sorts of datasets and problems. So like I said, it's available. It's Go. Of course, it would make sense if if we moved away from eBPF to use something like pcap if you wanna run it on Windows or Mac. And we also built a PIP installable module. So you just use that. You can pull in a pcap ML, and it will it it can pull directly into a panda's data frame and have all the columns for you. So with that, I'd be happy to take any questions.
[00:13:31] Session Chair: Thank you very much, Paul. So any questions? Please I mean, reminder, log in into the data tracker anyway as if you can record your participation. So Danish has a question. Right.
[00:13:44] Danish: Hi.
[00:13:47] Danish: Thanks for the nice talk. I wanna know what is the overhead of labeling while we are live capturing the packets.
[00:13:54] Paul Schmidt: So it we weren't able to sort of find a limit where it it fell on its face, although it was capturing on a a moderately powered desktop and on OpenWrt, and it it was a moderately powered OpenWrt router. So we didn't see the limit, but I'm certainly, it's not going to be free. And you're adding that overhead of storing these labels in your PCAP. You're definitely inflating
[00:14:17] Benedict: the PCAP size.
[00:14:18] Danish: Thanks. Thanks.
[00:14:20] Session Chair: Further questions? Maybe I have one in between. So you're actually using the common feed of the PCAP NG file syntax. Right? I'm just wondering whether this could also I mean, if you you only add five to metadata. I'm wondering whether it would help to ease and scale the processing if you would add additional data about the whole packet structure? Something like, oh, this is a TLS packet. Oh, this is a quick packet. Question is whether you check this in the processing chain of TCP dump and all of these filters, whereas the common field is processed before all of the details of the packet itself.
[00:15:01] Paul Schmidt: Yeah. That could make sense. So we kind of we we didn't do that at first because we figured these tools will be able to parse it. But, certainly, if you wanted to sort of offload some of that work, you could just put that directly in the comment field
[00:15:15] Session Chair: Yeah.
[00:15:15] Paul Schmidt: And make it make it lighter weight later. Yeah.
[00:15:18] Session Chair: Could could make the processing more scalable.
[00:15:20] Kilian: Yeah. Sure. Okay.
[00:15:21] Pei Jin: Yeah. That's a great idea.
[00:15:24] Session Chair: Cool. Any further questions? 123. No. Really cool tool. Thanks again. And yeah.
[00:15:32] Benedict: Thank you. So
[00:15:38] Session Chair: our next speaker is Benedict, who should hurry up. Quickly, quick, quick. He is a PhD student at Munich, and he will talk about NixNet NixNet reproducible virtual network experiments. So supporting measurements is the same here. Thanks.
[00:15:57] Benedict: Thank you for your nice introduction. So this talk is about Nixnet, reproducible virtual network experiments. This is joint work with Marcin, Paolo, and Jorg.
[00:16:09] Session Chair: Yes. Yep. Here.
[00:16:11] Benedict: Okay. Does
[00:16:15] Session Chair: it? It should work. Yeah. I gave you control.
[00:16:18] Benedict: This one doesn't.
[00:16:31] Session Chair: Yeah. Control. Let me stop.
[00:16:47] Benedict: I can also Can
[00:16:48] Session Chair: you Yeah. Just
[00:16:50] Benedict: the slides?
[00:16:50] Session Chair: Yeah. Just call it out to me. I'll change it. Oh. Yeah. I'll change it. Okay. Okay. Oh, good.
[00:16:55] Danish: Now it works.
[00:16:58] Session Chair: No. I I didn't. Yeah. Oh, okay. I I
[00:17:02] Benedict: just say next slide. Okay. Yeah. So networks network experiments are hard to reproduce in practice and harder than it should be. Next slide. Main problems are that dependencies of experiments are underspecified. Resources are no longer available. Package managers don't have the right version that you require to run it. The build environment that you are using is not the same that the original authors had. Same for running or or, yeah, for building although you might not know the CompileFlex that the authors used for running it. Similar thing. Different chip compilers, different interpreters. The host machine is different different OS, different configuration, and leftover applications from previous experiments or long running systems. Next slide. We want to be able to run today's network experiments tomorrow and even ten years from now, so we want to preserve them. Next slide. Why is this important? It's important across all the areas, academia, industry, teaching. We want to build on prior work and not reinvent the wheel over and over again. We want to test interoperability between different programs running at written in different languages, even run comparing old to new software and same or in in in some test scenarios, analyzing bugs, understanding their behavior in general. And, also, for teaching, it's important to, I mean, to have something you can share with students and not to that the students have to set up or spend a week to set up some test bed by themselves for a week, yeah, for a week open. Next slide. And what we are focusing on here is emulated test beds. So in contrast to simulated test beds, so simulation is more flexible but also more model based. Right? It's it's it's not testing the real software, but often we are interested in running real browsers, server applications in in an emulated testbed. On the other side, you can get more realistic results if you have hardware testbeds, but they I mean, they are expensive, time consuming to set up, so emulation is a good middle ground, I would say, and that's what we're focusing on in this paper and also talk. Next slide. The state of the art. There are many great tools already. You might be familiar with Mininet, ContainerLab, ContainerNet, or have handwritten your test setup with some batch scripting and TCNETDM. I mean, they are great tools, but they are facing some challenges. I mean, on the one on the one side, you have a lot of side effects from your system, either from your host system or also experiments that, yeah, leak in you to your experiment. Because dot files, configuration files, some caches that are placed somewhere, so no isolation, the dependencies are not managed. So your sys your your experiment is using programs that are on your host you might not be aware of, so some libraries, and you don't specify them because you don't know that your experiments actually depend on them. Handwriting is error prone. And if you use Docker's solution, mean, then you get a lot of isolation benefits, but they are more heavyweight. You have to use the images, and images Docker images are not reproducible from source per se, and they are quite heavy and also startup. Next slide. That's why we built NixNet. This is what we created and what the talk and paper is about, which manages the full dependency graph of the whole experiment. All the software is managed. There is no manual setup required. The build is hermetic, so isolated. The build is the same thing as your previous author did, and also the execution is isolated from the host system. So no leakage there. It runs on any Linux machine, on a single machine, so similar to a Mininet. It is you can run it as simply as just executing a sim single command. No root privilege privilege is required, which other tools often or mostly require, and no manual setup. So no leftoverstated can leak into your next experiments or fills up your disk. Just the results you care about. Slide. How does it work? Sounds good. Too good to be true. Next slide. Let's start with dependency management. Next slide. Next slide. Dependency is managed with many tools that you today use, like I mean, that's the state of the art. And the like, you version your soft software or code with Git. But, I mean, it's nice because it's tiny. It's only storing the diffs of code, but there's no instructions how to build this thing. And there is no instructions or no dependencies on which git repository is linked to others. I mean, you can hack it, but it's not built for that. Docker, on the other hand, is bundling all the stuff together so you can give it to someone. But giving it to someone is not the same as having the other or can recreate the same image because it's often relying on binaries that are fetched from some other dependency managers. And, yeah, like I said, no no tracking of the sources. Oh, what's going on? Okay. Docker images are quite large. I mean, if you want to preserve them for years, you have to store them. And they are in some registries, but they will be dropped after some time because either you have to pay or yeah, have a lot of disks around. On the other hand, there are those known package managers like a p g, mostly a really popular one. But this is a binary packet manager, so, also, you have binary packets that you fetch which are semantically versions. So there's a version number number on it, but this doesn't actually mean or you don't know how the the one who compiled this thing or the packet manager which which compile flags are used. Right? So if you have different Packet Managers, the version doesn't really mean anything that you get the same binary out of it. And also, if there are bug fixes, if there is version changes, the old sources or old binaries are dropped. And in network experiments, we actually want to reproduce or run experiments with the specific bug because, I mean, a year from now, the bug is fixed, but I don't care. I want to have the bug in my experiment. Right? Next slide. So
[00:24:20] Session Chair: would it
[00:24:21] Benedict: wouldn't it be nice to have something like a Merkle tree, like the the like Bitcoin or GitWorks that you have a dependency graph of a hash that identifies your your parts of the system And with a cryptographic hash, so your your your node is depending on the child node hashes, and if any of the child's change, your node changes, or your node has changes. And that's exactly what we want to have for experiments, a unique identifier with which changes when something dependency changes and stays the same. Next slide. And that's exactly what we are using. So this is NICS. So you might have heard of NICS OS. It's the exact thing it's it's built upon. So NICS handles all the dependencies of software that is used at runtime. So it's the runtime closure. So all the the the Merkle tree or Merkle graph of all the runtime dependencies, but also all the dependencies that are used during build time. So you have also built time graph with all the compiler source code, build flags, everything hashed in a nice Merkle tree. And, yes, that's what we use NICS for. Next slide. And with NixNet I mean, this slide is a bit complicated, but it shows that NixNet actually handles all of or on the top, you see the experiment. This is the hash of the experiment and all the dependencies. You can whatever you have in your experiment, configuration certificates, all the binaries, libraries, eBPF code, packet traces that you want to replay in your experiment. Like, think of anything that is a file, and this can be tracked with. Next slide. So one NICS expressions or NICS expressions are today used to build packets, but also operating systems, and now with NICS NET also network experiments. Next slide. Next slide.
[00:26:34] Kilian: Yeah. Okay. Yep. You have it.
[00:26:40] Benedict: Okay. This is just an example how it looks. So this is just a description with you describe the notes in your experiments. You describe your links. You describe the characteristics of the links. And, yeah, if you it's just JSON like. Next slide. I have hurry up, I think. And what Nixnet is actually it's just a really fancy bash script generator. But for a purpose, because I want to have low abstraction. So most of network engineers, we are familiar with how command line bash and or how you configure your firewall and NAT and stuff. So to be transparent, I want to have something readable or not not some Python syntax that's high on a higher abstraction level, so that's on purpose. Next slide. Now to running the experiments. I have to okay. I I could talk for two hours, but I I have to squeeze it in ten minutes. So we have built it, and, also, running is it also is also isolated. Next slide. We have a next slide. We have a multilayer isolation process. So we isolate experiment completely from the host system using the same technologies that Docker use. So this is Linux namespaces where you mostly use all the namespaces to isolate experiment from the system, but also the nodes inside. You have hundreds of nodes in your experiment. They are also isolated from this other because you don't want to have leakage between them. Next slide. This is what it looks. Just run mix run, and you get setting up this whole system, building all the dependency that you don't have built, running it. And in the end, you just get an output folder with all the artifacts you want to preserve. Other things is all cleaned up. Next slide. We test we evaluated this, not evaluated in a sense of we tested for the network performances, but we tested NixNet in terms of how fast is the setup. Because what we think is we want to have a fresh topology set up for every run of the experiment. So if you test because you do repetitions or you'd sweep across some large parameter sets, but you want to start with a fresh topology for every run. That's why I think setting up and clearing it down and using less memory is is quite important. Next slide. So you can see this blue line is the docker one. This is runtime. It's sort of startup and clear fill up phase. And we can for a 100 runs, eight nodes, we can save three dot five minutes just for setup and clear load running experiments, actually. Next slide. And it's similar for memory, peak memory. I mean, Docker containers, setting up the containers. It takes a lot of RAM, and we can do much better with Nixnet saving four hundred eight twenty eight nodes, 19 megabytes. There are some a lot of limitations. I think I'm over time already. Yeah. I can focus or, for example, macOS support is something we want to look at. I mean, we need we rely on network namespaces. That's why we are targeted to Linux already only, but there are we can use virtual machines. But because NICS runs on macOS already, so this would be some logical next step. But a lot of things there are GitHub issues for all of those Conclusion. Is I mean, NixNet is pinning the full source dependency graph of the experiment. It's lightweight using Linux namespaces. It's hermetic builds and hermetic runs, so no leakage between stuff. Low abstractions. Like I said, it's bash in the end. And so next experiment as well. We need to find our network experiments with next net. It's easy to share, preserve, and build upon. Thank you.
[00:30:27] Session Chair: Thank you very much, Benedik.
[00:30:35] Anthony: Maybe maybe we can get to
[00:30:36] Benedict: the next slide. Just the next slide.
[00:30:38] Session Chair: We are out of time.
[00:30:39] Session Chair: So
[00:30:39] Benedict: No. No. Just because there's a link on it. Just because if if you are interested.
[00:30:43] Shavit: So I really have a philosophical question. You Yes. Shavit from Tel Aviv University. You know, no man have ever stepped in the same river twice. If your experiment is so so sensitive to all these parameters, what is the meaning of the result?
[00:31:05] Benedict: Yeah. I mean, it's, yeah, it's it's nice to hear. I'm not sure if I could can answer that, but, yeah, it it's not deterministic. Right? It's not simulation. So, yeah, maybe we have to reproduce on purpose stuff rebuild stuff on purpose. But maybe if you want to have exact reproducibility or just build upon someone's work, it's probably better to have the thing at hand. But, yeah, not sure.
[00:31:35] Session Chair: Otherwise, you can continue the philosophical discussions during the break. There's more time.
[00:31:41] Audience Member: Hello. Thank you very much for your work, first of all. I would be interested in what does your tool offer beyond the pure dependency, management also, especially towards network impairments and mutation? Cost? That's what I'm
[00:31:56] Benedict: interested. Yeah. So for network impairments, I mean, we currently support I mean, technically, I haven't mentioned that, but, technically, you can also add your own hook points. So you can just call any Linux tooling. So our also, Nixnet relies currently on TCNETDM, so what Linux provides out of the box already. So we have also syntax for that to build it, but you can also write the bash line as you would in your terminal. But if you have your own emulator that you want to use, if you want to integrate NS three or some stuff in this, you can do that. But currently, we don't have syntax for that, but if feel free to open the issue issue or fateful request, so I'm happy to support more emulation tools.
[00:32:42] Session Chair: Thanks. Anthony, go
[00:32:43] Danish: ahead.
[00:32:43] Audience Member: Thanks. I'm in the same boat, so I'm already running it. You don't know how to sell it to me. So I'll give you one more good selling point. It's running different version of the OS. Like, if you want to have one side of the host with an old version and one side of the old new version, NICS is one of the best tools out there. So next, the real question. So are you using the NICS network topology builder, or are you using your own topology builder to create the networking topology you need?
[00:33:10] Benedict: What do you mean this next Topology Builder?
[00:33:12] Audience Member: So I'm assuming your tests have more than one Docker, or do you have more than one Docker?
[00:33:17] Benedict: Yep. It's not Docker, but, yeah, you have multiple network namespaces or multiple namespaces.
[00:33:22] Audience Member: So are you build connecting them already using Next Topology Builder, or did you build your own Topology Builder for that?
[00:33:28] Benedict: So we built a description, like a schema that generates the best batch code that creates virtual network cables with virtual network
[00:33:38] Session Chair: Yeah.
[00:33:38] Benedict: Yeah. Ethernet cables between your network namespaces, basically. So
[00:33:42] Audience Member: Because Nix Nix upstream Nix upstream has a thing, but that is very primitive. I couldn't build any complicated
[00:33:48] Benedict: I think then so NixOS doesn't I mean, NixOS can configure your whole network, but I they don't have certain tracks
[00:33:54] Audience Member: Nix testing. Not NixOS. Nix there's a Nix as a testing, and the testing is very hard to build the topology.
[00:34:01] Benedict: Okay. Okay. I I'm not I have not looked at the testing frame where TONICSOS is doing it. But okay. But that's a nice hint. Thank you.
[00:34:08] Session Chair: Thanks again. Thanks again to the speaker because we have to rush a little bit. Thank you. So our next speaker is Petia. She will explain explain about a Incident Triage
[00:34:28] Session Chair: I'll I'll run the slides for you. Measurements.
[00:34:32] Session Chair: And she is I couldn't fix it. Yeah. Sorry about this.
[00:34:38] Session Chair: We can try again if you want. I can try with the thing, and then if it doesn't work, let's do it. Hold on to it. Just give me a sec. Try it now.
[00:35:07] Session Chair: Try it now.
[00:35:09] Petia Vasilova: No.
[00:35:10] Session Chair: Then you have a human clicker.
[00:35:12] Session Chair: So if you have your laptop, I can give you slide control, and you can present from your laptop. It's easier for you.
[00:35:19] Petia Vasilova: Can I stay next to you?
[00:35:20] Session Chair: Yeah. Go ahead.
[00:35:23] Pei Jin: Thank you. Welcome.
[00:35:25] Session Chair: I don't if the camera's gonna
[00:35:27] Sam: track you.
[00:35:28] Petia Vasilova: Okay. Thank you.
[00:35:30] Session Chair: Okay. Let me take back control from the clicker. Okay. Now it's yours. Just right arrow, left arrow.
[00:35:37] Petia Vasilova: Yeah. Okay. So hi, everyone. My name is Petty Vasilova, and today, I will talk about the explainable routing aware incident triage from active measurements that we implemented for the needs of the WLCG network. So
[00:35:58] Danish: Just give me a sec.
[00:36:01] Petia Vasilova: Space? No?
[00:36:05] Session Chair: Right there. Yeah.
[00:36:07] Petia Vasilova: Okay. So LHC stands for Large Hadron Collider. This is the biggest machine ever built. Its purpose is to produce subatomic particles and smash them in in the collider. And the reason for doing that is to reproduce the conditions, right after the big bang. So particle physicists collect this data and analyze it, record it, replicate it, and so on through the, worldwide LHC computing grid. So the home for this collider is CERN in Geneva, Switzerland. And the the about 20% of the computing and storage capacity is really at CERN. The rest is distributed across the world. Thus, needs to flow between sites continuously. The last reports I read was saying that scientists transfer up to 10 petabytes daily of data. So a reliable network and high performing network is crucial for the scientific community. Thus, we deployed personnel toolkits. I don't have time to explain why what personnel does here, but what we do is we actively measure the network, the network state, and we record. Thank you. Okay. Nice. And we we stored the recorded data in a central place in Elasticsearch database at University of Chicago. We test different aspects of the network such as latency throughput and the network path through a standard trace routes. And we face many challenges. The motivation for this work is that we wanted to use all available data and correlate somehow the network path with other performance metrics. So some of the challenges we face are, for example, the noisy trace routes, the incomplete tests. Also, have this challenge that the way we build the meshes is that they run the different tests run at different cadences. For example, the latencies are once every minute, the trace routes once every ten minutes, while the throughput is much heavier tests, so it's once every six up to twenty four hours. You can imagine how hard it is to align those measurements and correlate them with each other. So whenever we have a complaint that there is a slow data transfer, we would like to know first thing we would like to know is whether the path was usual or somehow anomalous in order to see where the problem started and what is the location of it. So we came up with this idea of creating this framework, Routing Aware Incident Triage framework. It could be split in two lanes. The first one is processing only the trace routes. Essentially, we built a statistical model that can tell us whether a path is anomalous or normal. And the second lane is dealing the with the performance metrics, the one way delay, packet loss, and throughput. This is much simpler because we just use thresholds derived from experience and from past data. Finally, we align those anomalous data sets in time and produce an outcome triage where each category is not only descriptive but gives us priority for the people investigating network issues. This is a plot to show you that we tested, we validated our framework on three historical cases. One is the Singapore hardware failure. Then we have one recurring much sub store case in Spain, and the planned migration of a single router. All of these have different signatures that we can see in our data. And what I wanted to show you here, if we take the Singapore as an example, we started from about nine 4,000 pairs, monitored pairs, and through different quality gates and triage labels, we managed to focus on only about 120 monitor pairs that we need to pay attention to that were affected by this incident. So how do we represent the paths? Naturally, we build a graph where each node on the graph is either an AS number or a slash 64 prefix. The edges are could be a membership or and at the same time, we assign weights to them that are representing the real the real usage of those links. And also, we we bias the walks that are something we we create a set of 300 weighted walks, And they are biased by source s because we want to keep our graph representation as much as possible close to reality. At the end of this step, what we essentially have is a note embedding. For each of the notes, we have a vector of 32 dimensions. And the intuition behind this is that nodes that go together very often, they will cluster together, while the new nodes that are unusual will be measurably further from the cluster. Okay. Then we want to know whether a single path is anomalous or not. So what we do is we take all these embeddings for each of the nodes. And for each path, we average the node embeddings that are seen on the path. In addition to that, we build the structural features, which is which are three features that represent the IP IP repetition ratio. So whether the path has a lot of loops or not, TTL gap increase ratio, and path length, which is essentially whether the path became much shorter or longer than expected. In the end, we combined the two embedding and structural values into a single composite score. We don't have the time for that, so I'll jump into the examples. So this is a life this is an example from our life application from the Singapore hardware problem. Those four plots represent each of the components. This one is about the composite score. So you can see that initially, the score was about 0.3. And suddenly, when the issue started, it jumped to the maximum of one. Then the second plot is the one way delay, and this is the interesting part because if we see here the normal for this pair, which is Southwest tier two to Taiwan, was about two hundred twenty milliseconds, and suddenly it became much lower, about eighty milliseconds. But when we look at the packet loss, we see a significant increase, and the throughput is close to zero. If we look at two examples of comparing paths from the baseline dataset and the the anomalous dataset, We see why this happened, and initially, the baseline path was using the ESnet, I believe. So this is the expected behavior. It should go through the research identification network. And then in the anomalous path, we see commercial transit through other AS numbers. So detector that was based entirely on delay would consider this as improvement. But since we use all available measurements, we can say that there was a problem and this pair. Finally, this is a live capture. In May, We saw this problem in our application first. It was a submarine cable fault between UK and Netherlands. This is a heat this is a heat map from our application, and the columns represent the destinations, the rows are the sources, and the colors are the different triage categories. So we can see that The Netherlands column and The UK row highlighted in this plot, and the NOC ticket confirming the incident came hours later. Okay. Conclusions. So using this method and this workflow, we can not only say if something was wrong, but we can actually pinpoint what changed, where, and how urgent it is to investigate. I should say here that we do not use any other data source except for the personnel measurements. No BGP feeds. No additional probing. Only the tests we run with personnel. And this is the URL to our app, which is still a development version. So it may be a little bit slow, but you can check it out. Yes. So this work was supported by the National Science Foundation through the OSG, Sun Project, Iris HEP. And I think yes. Thank you.
[00:47:29] Session Chair: Thank you very much. Go ahead. Next question.
[00:47:38] Session Chair: Hi. Nice talk. Anant here from Netflix. You probably have a lot more latency data than trace routes. Right? So was there a granularity level that you thought worked better to be able to correlate those two datasets?
[00:47:52] Petia Vasilova: I cannot hear you very well.
[00:47:54] Session Chair: You probably have more latency data than trace routes.
[00:47:58] Pei Jin: Yes.
[00:47:58] Session Chair: Yeah. So was there a granularity of trace routes that you think work better to connect those two? Because you wanna stitch them together. Right? Like, what frequency were you collecting the trace routes at?
[00:48:11] Petia Vasilova: So the trace routes are once every ten minutes.
[00:48:13] Session Chair: Once every ten minutes? Yes. Was that enough to capture any short duration events?
[00:48:19] Petia Vasilova: I I would say not enough, but this is what is allowed because it takes a lot of space, all this data. We received about 4,000,000 tests, and this is even half of what we used to to get. We should have even more data than than that.
[00:48:35] Pei Jin: Okay.
[00:48:35] Petia Vasilova: So I I tried to convince my colleagues to increase the rate, but this is as much as I could get. And because this is once every ten minutes, I realized that I shouldn't be focusing on short lived problems
[00:48:55] Session Chair: Mhmm.
[00:48:55] Petia Vasilova: And focused only on the big problems Okay. Like that that they last longer.
[00:49:02] Session Chair: Sounds good. And one quick follow-up. You are making your notes at AS level and then associating the prefixes. Were you also interested in looking at link level because you get the hops? Right? So are you doing any link level analysis?
[00:49:15] Petia Vasilova: I I I couldn't hear you.
[00:49:16] Session Chair: When you get the trace routes, are you doing any link level analysis in them, like, per hop, or are you collapsing it at an AS level? So if you had five hops within an AS, are you putting it as a one node, or are you doing any, like, basically, five links within that AS?
[00:49:31] Petia Vasilova: So we use the information about the ASSs and also the we use the IP addresses, but I course them to slash sixty four. And we can talk later.
[00:49:45] Anthony: I'll pull it up.
[00:49:45] Petia Vasilova: I have a lot to talk about.
[00:49:47] Benedict: That's good. Thank you.
[00:49:48] Petia Vasilova: Yes.
[00:49:49] Session Chair: Good. Thanks. Other question? Okay. Go
[00:49:52] Sam: ahead. Hello.
[00:49:53] Petia Vasilova: Yes.
[00:49:53] Sam: I'm Sam from University of Trenton. I have a question regarding your active measurement tool. So there are already publicly available tools like Ripe Atlas, Cager provides those probes. Why, instead of using that, is there any specific reason of using your own probe?
[00:50:13] Petia Vasilova: I cannot really answer this question because these decisions were taken way before I came to work with my colleagues. But I know that the project is built with with some of the participants were in my group, and it was a decision on the level of the community. I can check this for you.
[00:50:41] Session Chair: Okay. Thanks again, Petra.
[00:50:49] Danish: Just a quick note in between. ACM just informed that the papers are online now. We are about to put them also on the website with the links, but they're in the ACM digital library.
[00:51:02] Session Chair: Maybe you can post a link in the chat of Viteko to the ACM digital library because they are open access anyway. So our next presenter is Antony. He's an almost first year PhD student with Olivier at UC LeRonde. And he will talk about some are old, but still a very important topic that becomes even more important considering the current traffic. It's about interdomain multicast routing for modern routers. Please go ahead.
[00:51:31] Anthony: So hello, everyone. I'm Anthony. And today, I'll present you present you our paper. And I will present you our paper about multi cast routing for modern routers. So before starting, let us review a bit how multi cast works and basically how the PIM routing protocol works. So, basically, multicast is a point to multipoint communication scheme where one source send the same data to a set of receivers. Basically, a client will notify that they want to receive some multicast data from the source by sending an iGMP or MLD join to the towards the source. And it basically indicates that you want to receive some data from a source s, which is I'd identified by a specific group g. Upon receiving this join, the gateway will forward the join up to the source by using the short sparse. So, basically, here, here, four sends the join to r foo r two, which forwards it to r one. And, basically, the multicast branch is created. Then if we add a second client, because when we are in multicast, we want to have multiple receivers. We can do the same for client two, which once again is an SGMP or MLD join, which basically depend on the IP version, which will be forwarded by a father on the shortest path to the server, which is via r three, then via arrow, which itself forward it okay. Then one okay. Which forward it to r one. And then we have the multicast fee which is created and is highlighted on the slide in green. So every packet multicast packets sent by the server will follow the different branches, and we arrive at the the different clients. So to create this multicast tree in the top lane, we need to have some state that are stored on the routers, which is derived from the joints. And, basically, as r one received a joint from r two and r three, it's known that for a specific topple, s g topple, it should replicate the packets and send them to both r two and r three. So one of the problem with multicast is that you need to off state. And every s g states, which correspond to a multicast flow, require an entry in the multicast FIB. If we look at the hardware amplitude capacities of different routers, we can see that they span multiple orders of magnitude ranging from 1,000 entries for low end devices up to several hundreds of entries for high end routers like Cisco 8,000. So what could happen in practice is that you have state exhaustion. And we highlight this in our revised example. So let's consider what that we know of limited number of multicast fib entries on routers, and we add this to the example by showing the maximum number of m fib entries and also the remaining number of m fib entries for each router. Basically, what PIM does is that if we rerun the example, we will send one again once again a join. And upon receiving the join, we install the entry in f m fib, so we consume one entry on f four. Then forwards it regardless of the m fib entries to r two because it doesn't consider these entries. And and r two sees that it has no remaining m fib entries. So what it does currently is simply it ignores a multicast join. It does not install an additional entry. And this is quite annoying for the client because it joined the multicast group, but in the end, it does not receive any any multicast data ever. And from the client point of view, it is even annoying. It's even more annoying because it doesn't really know the reason for this lack of reception. It could be due to exhaustive states. It could be due to multicast configuration issues at the routers, or it could even be a packet loss. And depending on the actual reason for this pro for the the lack of reception, the actual reaction that the client should take is not the same. So this is quite annoying. So in our paper, what we propose is to expose these multicast capacities in the control plane and then use this information to roll the joint around rotors that have no remaining entries. This is the first part. We want to expose the entries. And to to do this, we simply reuse the existing link state protocols. So basically, we use OSPF or ISS and add an additional TLV that includes the multicast capacity and the remaining number of entries. And then we rely on the LixTEM protocols to flow this information in the old network. And eventually, thanks to flooding, every router has a global view of the remaining m fib entries for all other routers in the network. In practice, we don't really want to flood this information every time an entry is installed. So in the paper, we propose a threshold based approach with hysteresis, which basically means that we send the info new LSP with this capacity information only when the number of entries are significantly significantly enough changed. So this is quite similar to the traffic engineering exist extensions that are already existing for SPF and ISS. So once we have this information, we need to find a path because we want to avoid these exhausted routers. And to achieve this, we propose two path computation approaches, which are basically pruned shortest path, which simply remove the routers with exited entries from the graph and then run a simple shortest path on the resulting graph. Or we can also use the shortest widest path, which basically try to find a path with maximum number maximum number of remaining entries along the path. So in both cases, on the example, we end up with the green path where we want to use the path via r five and r three, then r one, which has plenty remaining entries to install the multicast state. In practice, when we want to we send the join, we also encode the full path in the pin join simply to avoid defect deflections that result from delayed LSPs or diverging LSDBs. So we encode so we put the full path in the pin join, and then the pin join is simply sent over the green path. So that's nice. But how does it work how does it performs in practice? To do this, we did simulations on three different topologies from Topology Zoo and also on FATRI topology where we measured a joint acceptance rate, which is basically simply the proportion of joints that have succeeded, which means that they have reached the router directly connected to a source. And we measured this acceptance rate according to the varying network load. So here, we define the network load as the mean number of concurrent multicast trees that are present in the network. And we measure just acceptance rate and connect to it. In practice, during the simulation, we use random random number of multicast entries in the range 55,000 to represent the spread that we have seen on the real routers. So what we can see is that at low loads, we don't have any difference in the joint acceptance rate, which is totally normal because the network has plenty remaining entries, and then they are never saturated. But if we increase a bit in the load, we can see that it becomes more interesting as our methods, both Grund SP and the shortest YDIF path, allows to have roughly 2022% higher join acceptance rate compared to classical PIM, which directly translates to more concurrent if it wants to go to next one, which directly translates to more concurrent multicast tree present in the network for the same routers. So we have the same network, but we can have more concurrent multicast trees. And if we look at the other end of the spectrum, we can see that high load, we don't see any difference again simply because the network is saturated, and we don't have any alternate path to route the joins on. So that's nice. We can have higher join acceptance higher join acceptance rates, but every gain comes with a trade off. And, basically, the solution trades the path length optimality for a higher acceptance rate. Because when we erode the joint alternate path, they may be longer to accommodate the multicast FIB capacities limits. But overall, we can see that the overhead is roughly limited because we are roughly between one and three additional ops for all the for all the multicast joints. In practice, you also have some impact on the flooding because you need to send this information through the network to inform the other routers of the capacities. And so you need to flood this information and keep it fresh. One interesting observation is that when we increase the network load, the flooding is more impacted, which is, in fact, directly represented by the the fact that when you have more multicast flows, you have you use more alternate paths. And if these alternate paths are also overloaded, Basically, they need to send more information to say that the multicast entries have been exhausted. And finally, you can also see that depending on the actual topology, the number of advertisement changes simply because the number of alternate paths that are existing in these topologies may differ. Great. So in conclusion, in our work, we have shown that PIM does not consider the multicast capacity currently, which can directly lead to multicast tree creation failures, which are quite annoying for clients. In our paper, we have proposed to expose these capacities with existing links of protocols such as OSPF and ISS, and then use this information to route the joints around exhausted routers. Using this approach, we can increase by roughly 22% the joint exponent rate, which directly translate to more concurrent multicastries for a similar network. As future work, we are also considering how we could use other metrics such as bandwidth and delays in the joint creation in the tree creation process such that we could have bandwidth optimized or delay optimized trees. So thank you for your attention. If you have any questions, feel free to ask them.
[01:03:03] Session Chair: Thank you very much. So let's start with questions from Max.
[01:03:12] Max: Hi, Max. Nice work. One question about considering that multicast, often you know the capacities of the flows, right, and they're predetermined and fixed. So for video stream, it would be a certain bit rate. Have you also considered including the available link capacity in your in your idea to make sure that the links between the routers also have support, the capacity that the multicast stream would require?
[01:03:38] Anthony: If I understood correctly the question, so you asked if you could include link bandwidth in the tree creation process. That's it.
[01:03:45] Max: Basically, yes. So you can, at the same time, make sure that you route the multicast in a way that the link is supported.
[01:03:49] Anthony: Yeah. You you could do. We don't do it in the paper, but, basically, you have the you assume that you have links to protocol, so you can do whatever you want on the join, on the path computation for the join. So if you want to have a multicast read that you know will use, like, 16 megabytes, megabit per second, you can say, okay. I want to exclude all the links that have only eight megabit per second capacity. So you can reroute the joins on these paths. Okay.
[01:04:14] Max: That might be a worthwhile extension.
[01:04:16] Benedict: Thanks. Thank you
[01:04:17] Anthony: for your question. Thomas?
[01:04:21] Danish: Some protocol engineering questions or remarks. PIM is called protocol independent multicast, and that means it's independent of the unicast routing protocol. So, I wonder, could this be engineered in PIM itself, not using OSPF or ISIS? That would be the first question. The second question would be a bit I mean, the whole signalling gets only interesting once you're approaching overload. Typically, signalling will also lead to rerouting that is sort of local, typically. The question is, wouldn't it be more appropriate to only signal, instead of flooding all the time, the stuff only signals overload or close to overload thresholds to a local region so that you can circumvent the problem of overload you you address?
[01:05:19] Anthony: Thank you for your questions. So basically, indeed, PIM was protocol independent. But if you want to have this information, capacity information, you would need to have info information about the old topology. And what we did was assume that we have a link set protocol because it is basically does basically become the norm. And we inspired ourselves from segment routing, which also try to do MPLS by assuming a link state protocol. So we took a similar approach than segment routing. And for the second question, which was about
[01:05:55] Danish: Local. Local threshold.
[01:05:58] Anthony: This was in the backup slide, it's the only so basically, the thresholds are not balanced. So we try to put coarser bands at low utilization. So, basically, you have a threshold at 50% utilization, then at 50 75, then 87, which makes that the threshold bands are now at high utilization. So that if you don't if you have plenty available entries, you don't fill the network always. So we would try to take this fact into consideration during the the proposition.
[01:06:39] Session Chair: Okay. Thank you very much again. And our next speaker is Faribhar Shi is with the Maxbank Institute in Saabrucken, and she will talk about how to detect I p v four subnets in the wild and why it is important.
[01:07:00] Session Chair: Yeah. Just like next step. It's good.
[01:07:06] Faribhar: Good afternoon, everyone. I'm today, I'm
[01:07:10] Session Chair: Mike. Okay. Sorry.
[01:07:15] Faribhar: Good afternoon, everyone. Today, I'm gonna talk about our, work detecting IPv four subnet in the wild. And this is joint four with MaxSpline, Tuberle, and Tiedel. So why is measuring subnet is important? BGP announces prefixes, and this is the aggregated view of the the the information. However, network operators inside of their network, they partition their prefixes into subnet, a smaller network. And this subnets usually a structure of the subnet that they are using internally is not feasible for us. And this is exactly what we want to identify or discover in this work. This may be a help us to better understand that the better understand help us understand that IP resource utilization or the network operation task and also will give us more accurate abuse handling, block block listing, geolocation, which is published in the RFC recently published in the RFC 1997. Before going in the detail, I'm going to give some overview of our methodology. We used two complementary methodologies, which is one of them is the ICMP echo, and another is the ICMP address mask. We all know that each subnet has a network address, broadcast address, and subnet mask, which defines the boundary of the subnet. For echo methodology, we use we ping broadcast and network address, and based on the responses that we get, we identify the we we somehow tell that we infer the structure of the the subnet. And I'm gonna first focus on the ICMP echo, and then later I will explain the address mask as well. How does it work? In we have over like, in this simple example, house and the router over there, and we want to ping the router. In the normal situations and the norm that we expect, it reply the rotor replies with the switching the source and address destination like the simply reverse. This is what what is exactly explained in the RFC nine seven hundred and ninety two, just switching the destination address and the source address in the in the just quoted here as well. Now, unfortunately or fortunately for us, it's not always the case. Sometimes we ping the broadcast address like this, and the answer that we get is different from the that supposed to be. So this this is the help us to do our measurement here. There are some legitimate reason for using that different address for the reply, which is like there is maybe in the different configuration, a specific configuration in the routers, or the load will answer, or look back address answer the request. But what we are interested in this work is to see how different routers handling broadcast broadcast and network address ping. To do so, we have a very simple setup, like simple experiment lab in our inside network, which is controlled environment, and we test a few hardware available that we have in the experiment. This is the schematic of the one that we did in our lab experiment. We want to probe the network over here, and every time we check the other we replace the rotor in the middle with different hardwares. By sending the echo request for example, in the sorry. In the Juniper case, by sending the echo request for the network or broadcast address, you were simply by response with the address that PIN sent. Oops. For the MitraTik, it simply dropped the request. And for Cisco, more interestingly, it differ it behaves differently, and it puts the the receiving interface in the response to back to the host. And we know that different we cannot generalize it that because because we just test a few hardware in the lab experience, so different OS version, they might behave differently. But it can allow us to reveal subnet boundary in the world. So did this experiment, and we used this method in a while. We crafted ICMP messages and put the destination IP in the payload. And we did all measurement measurement, our measurement, and all routable IPv4 addresses on November 2025. Then, if the expected behavior is like the putting in the response, the payload and source are there would be same. If there's a mismatch, we mark them and filter them out for the subnet detection. Based on this method, we identified 5,000,000.2 addresses, which we detected 2,000,000 subnets. But to see how different vantage points have effects on the R measurement, because every packet in the Internet travels different routes, and they have a different path discovery, we have four vantage points, one in The US, two in Germany, and one in Osriola. And to see that our detected subnets, how they look like, this offset plot shows that overlap of the different vantage point detect the subnet in the different side ventral, and the bottom, you see that different combination of those vantage points. And the bar shows that how many subnet that we take in each combination. 55% of the detected subnets we discovered from the old vantage point. 14 of them, they are only visible in the one vantage point. And 31% is discovered by two or three. So we can say that having more VantagePoint can greatly affect the ICMP echo method memory. We just checked with the four, but we left that how many VantagePoints we need as a future work. There's another element in the cool element in ICMP, which is ICMP address mask. Simply, a half wants to learn about the next mask, put the zero in the mask, and the gateway answer that with the sublet mask in the the reply. This is deprecated. I simply message which was introduced in with RFC 950. It is deprecated. But it doesn't mean that it's not it's it's not supported actively, but it doesn't mean that they aren't used in the deployed in the Internet. So we are using that to find the subnets as well. By this method, we find 2,300,000 addresses, which gives us the 1,800,000 subnets. And the these two method contribution is over here in this plot, and we've overall defined less than 20 minimal 25 k overlap on both method because I think the address mask by default disabled in Cisco devices, but there are other vendors that are using this. They are still using address mask, and they are enabled in their devices. And, yeah, as you can see that in the plot, SH30 and SH31, they are most common in our data set. But we're also interested to see the pattern of the vendor that we see in the data as well. We did the SNMP version fingerprinting technique, which is introduced by Taha et al in IMC21. And among those addresses that we find as a mismatch in the ECHO method, 44% of them, they respond to ICMP version three. And as you can see in this plot, 98% of them, they were Cisco devices. And there are two percent margin of the error, we say, because we cannot confirm their behavior. In the second method, we only found 25% of those the vendor of the 27% of those addresses. And this is a top five vendor that we can see here, which is responsible for the 90% of these 27 responsive addresses. So based on the ICMP echo method, which is based on the data we find in our measurement, is only applicable in the Cisco devices. And also, on the lab experiment that we did, it indirectly validates our finding. But in contrast, several vendors allow ICMP address mass replies by default, and they allow that over the Internet. To summarize, based on these two methods that we proposed, we find 3,800,000 of the subnet mask. And we want to look at later on to see what we can find more in that address set of the future. Thank you.
[01:18:41] Session Chair: Thank you very much.
[01:18:48] Session Chair: Yeah. I do have a question. So the directed broadcast in ICMP have been disabled in a lot of Cisco devices for a very the directed broadcast.
[01:19:01] Faribhar: I can't hear you properly because
[01:19:03] Session Chair: Can you hear me in the room or
[01:19:05] Session Chair: is this Very well. Okay.
[01:19:07] Session Chair: If you can repeat what I said. Right? Like, the directed broadcast have been disabled in Cisco orders for, like, a very long time. So what is the configuration you tested with?
[01:19:15] Faribhar: Oh, direct broadcast is the the configuration for the address address mass in the Cisco device. And and by default, because of the some abuse and patient that they have, it's by default disabled. By default disabled.
[01:19:31] Session Chair: Correct. So you just you enabled it to test it. Is that
[01:19:33] Faribhar: For the Cisco device, it no. Okay. Because we just measured over oh.
[01:19:40] Session Chair: But maybe then the question is which firmware version did you consider? How old was the
[01:19:44] Faribhar: firmware Oh, yeah. We have a few devices in there. For Cisco, we have 2,900 series. And for Juniper, we have MX 150. Mhmm. And the microtether is it instance of the the post.
[01:20:03] Session Chair: K. Thank you.
[01:20:04] Sam: Any further? Interesting study. I'm Sam from University of Georgia. Sorry if I misunderstood something. For example, in your finding, saw dominant of dominance of slash 31, but that subnet doesn't have broadcast address, and I think your method relies on broadcast address. How is that possible? Another question is, did you validate your finding with network operators?
[01:20:31] Session Chair: Sorry. Did you Oh, you you you didn't hear?
[01:20:33] Danish: Find with network operators.
[01:20:37] Faribhar: No. We just tested ourself in the the.
[01:20:41] Sam: And about the first questions, since your method relies on broadcast address and slash 31 doesn't have a broadcast address, how did you detect that?
[01:20:52] Faribhar: We did some filtering. The study is already gone. We did some filtering if the address broadcast and the network address, they respond by the same same address. And in the middle, because there are there might be other servers in the that they response at, we keep those only that's that's as much as other than that, we dropped it. It's it's not there anymore. I can show it later in in my Yeah. Result.
[01:21:20] Sam: Maybe I'll read the paper. Thanks.
[01:21:22] Petia Vasilova: No problem.
[01:21:23] Session Chair: Then thank you very much. With the next speaker who is Kirian, he's a PhD student in Georg Karlskrupp at Munich, and he will talk about reliable VPNs with mask.
[01:21:40] Kilian: Yes. Hello. My name is Kilian. Recently, we did some investigation on reliability improvements of virtual private networks, and I want to share some results of our investigations. Typically, VPNs are used for establishing end to end connectivity between devices and to encrypt the traffic along the way. However, there are some more demanding use cases that demand reliable properties of VPNs, for example, for legacy applications, provide area networks, or also for site to site VPNs. And when we talk about reliable VPN properties, we mean loss recovery, so low loss, and minimum delay. There are some basically standard software around that are frequently used. One example is WireGuard, which is part of Linux. It uses UDP for transport, so it's unreliable and has a rather minimalistic feature set and only operates at layer three. OpenVPN, in contrast to that, is a user space program, offers reliable transport over TCP and offers unreliable transport as well over UDP and can configure can be configured over layer two and layer three. With mask, we have a new, basically, standardization effort which builds on top of HTTP. You typically implement it in user space. We can use HTTP capsules for reliable transport or HTTP three datagrams that then can map to quick datagrams for unreliable transport. And now we have basically standardization activities at layer two, three, and four. So what mechanisms do we have available for building reliable VPNs? Retransmissions are basically well known. They are part of TCP and used in quick streams. We need to be careful there. TCP meltdown is a problem that occurs when we nest different congestion controlled protocols that use retransmissions within each other. And the congestion controller gets confused by increasing delay that is basically coming from the retransmission so the throughput can can sink. So, yeah, so we need to be careful about this problem. Another opportunity that we have is to use forward error correction. Recently, there was a lot of research activity there that discussed different extension of or experimental extensions of QUIC. So there's not a standardization standardized extension yet, but many things are moving there. And also, FortiGate has a VPN appliance, which uses some block coding scheme. We can get some independence of problems on an individual path by adding additional path using multipath connections. There are experiments with WireGuard and OpenVPN in this regard. But also now, we also almost have the multipath quick standard, So we can use that to build a multipath mask connection. And we can go even further by then building a network overlay by nesting different mask session within each other. It's depicted in the figure on the bottom here. We have a client and a server and the autonomous system a and d, left and right. And then we basically use an VPN relay, an autonomous system e, as a jump host. And with that, we can, for example, circumvent peering issues by basically performing some sort of source routing or generating synthetic path with it. To perform our investigations, we built a special measurement setup We're using NetM for the path property emulation in a symmetric configuration. And client on server side, we use the different VPN software, which is WireGuard, OpenVPN, and Mask in different configurations. For Mask, we built our own prototype, which is based on the multipath Cloudflare Quiche work by Quentin Deconink. And then we extended it it with Tetris World error correction scheme and connection nesting. For the scenarios, we use different properties. For example, the normal property only uses symmetric twenty milliseconds of delay. And then we have scenarios, for example, for random loss and burst loss where where we use a Gilbert Elliot loss model. The measurement traffic is generated by iPerf three in UDP mode. We use the really slow measurement rate of one megabit. So why are we using such a low measure measurement rate? Initially, we measured the throughput without any impairments. And there, we achieved with FireGuard and OpenVPN a throughput of around 1.6 gigabit per second. Also, it's not worth it to see that there is a bit different behavior of WireGuard and OpenVPN. With WireGuard, apparently, we weren't able to basically put more traffic inside the WireGuard socket. With OpenVPN, we could put more into the TAN socket. So it's behaving a bit differently there, but still the throughput was one dot six gigabit per second. With mask, we achieved 600 megabit per second. We didn't do any further optimizations, but for our measurements, it was fine. We are living with this rate. And also, there are other reports of using CloudFacush with much higher higher rates. So I think there's opportunity for improvement. And then we switched on random and burst loss and repeated those measurements and with mask over datagrams. So the unreliable mask case, we already encountered issues at the slow rate of one megabit per second. And then we did some manual investigations and figured out that congestion control actually is blocking us because it treats those lost packets as congestion signals and somehow mischarges basically the available bandwidth. So it's basically an advantage and a disadvantage having congestion control within the virtual driver network. And then we went ahead and did the actual delay and loss measurements at at this rate. Initially, we didn't configure any reliability. And already here, we see some differences between WireGuard OpenVPN and Mask in the and the, basically, UDP or datagram mode. We we see a bit of additional overhead or a bit of additional delay for OpenVPN and mask because they are user space user space programs in contrast to Viagra, which shows a bit less delay in this case. If you switch towards the random and burst loss measurements, we see that mask shows some higher tail latency here. We can again attribute this to the issues with the congestion control that we encountered also in the previous measurements. And we see also that the loss percentages in the table on the end to end path and the inner path are a bit different. This comes from handshake processes, handshake packets, and additional signaling packets being exchanged. Next, we switch to a reliable transmission. We see that on the end to end path, there are no losses. And we also noticed that our VPN somehow introduces some additional delay in comparison to mask. We didn't investigate this further, but it's still interesting find. And as soon as we switch on the random and burst loss, of course, we get a high tail latency, which is needed for the retransmissions of around one hundred milliseconds. On the end to end loss, we don't have any on the end to end path, we don't have any losses. However, on the inner inner path, of course, we have losses that are corrected by the retransmissions. If you use forward error correction, we are basically down to only having our mask implementation available. They are the end to end to end latency shrinks or one way delay shrinks to forty four millisecond. Of course, this depends on the proper tuning of the forward error correction, sending enough repair symbols to be able to reconstruct it. So it requires some kind of previous knowledge on the on the link characteristics. But in this case, so this was our our result that we could measure. And also, there is no end to end loss visible here. Then we extended the setup, adding a second link between client and server, and performed the multipath measurements. We kept one path with the normal configuration, so the twenty milliseconds of symmetric delay, and iterated for the second pass through all other impairments that we showed in the table previously. We found out that on this bottleneck and high delay scenario, there was an increased tail latency. For one, it takes some time for the quick implementation to gather the path metrics and before it switches to the other path. Also, the multipath scheduler somehow becomes a critical element of our VPN system now. So we need to be careful which scheduler we choose. In this case, we use the low RTT scheduler, which basically selects the path at the lowest RTT so it doesn't react, for example, to losses. So we also have some loss on the end to end links here. And then finally, we extend the setup further and add basically this nested connections with a network overlay, generating basically artificial path in between. And in this case, the scheduler basically chooses the perfect path, and we get with our configuration here a much lower delay, giving us the opportunity basically to improve the end to end link quality. So in conclusion, mask, the congestion control is always on in comparison to all other VPN systems. So we can, if we are not careful, can perceive blocked forwarding on lossy parts. Our prototype currently achieves less throughput at OpenVPN and WireGuard. We evaluated mask with retransmissions, forward air correction, multipath, and overlay. With the exception of forward air correction, all of those mechanisms are standardized or nearly standardized, let's say. And we found out that forward error correction and retransmissions both avoid end to end losses with the advantage of forward error correction that we have the opportunity to reduce the delay that is needed for loss recovery. And with multipath and overlay, we have interesting tools available to circumvent issues of low quality parts. And we have some more results in our paper with with high delay measurements, bottleneck measurements. And that's it. Thanks for listening, and please ask any questions.
[01:32:55] Session Chair: Thank you very much. Any questions? And maybe a brief question regarding the implementation. Is it available for which platforms?
[01:33:16] Kilian: Linux, Windows? No problem. It's a Cloudflare Quiche library. It runs really on Linux. I think it also runs on I think it should run on any platform since it's basically used as Cloudflare's VPN gateway, which can be installed by anybody. So I think it's quite a. And
[01:33:36] Session Chair: how much I mean, do your results depend on the actual quick library? I mean, is there any side effects that might occur?
[01:33:46] Kilian: So this forward air collection library is kind of independent. Of course, you need to hook it into the proper parts in the QUIC library. Mean, you need multi pass support, which is now coming into towards more QUIC libraries. For the connection nesting, I think this can be done basically as part of the application. So most of it is available. And it would be interesting to also have the void air correction standardized somewhat to make it usable for more people.
[01:34:15] Session Chair: Okay. Cool input with ITF. If ah, one more question?
[01:34:18] Audience Member: Yes. Yes. So, yeah, my apologies if you covered this already at some point. But
[01:34:23] Sam: Please state your name. What
[01:34:25] Audience Member: are some use cases where Oh, your name? I'm I'm from Chinese New York Hong Kong. And, yeah, apologies for if you covered this already. But what are some use cases where we want VPN to be reliable? Right?
[01:34:40] Kilian: Like Mhmm. Yeah. Yeah. So they are really old applications that are not intended to be operated over wide area networks, but somehow they need to be operated over wide area networks. So this is one use case. But I think also in general, side to side VPNs, you can benefit from such improvements, I think, to have more robust connectivity between them.
[01:35:06] Session Chair: Okay. Thanks. Other question? Nope. Then thanks again. So where's our next speaker?
[01:35:20] Session Chair: Pei Jin? Pei Jin. If you're remote, like, ask me for slide sharing permission, I'll grant it.
[01:35:25] Session Chair: Oh, okay. So Pei Jin, he's with University of Maryland, and he will talk about RTC Relay server relay server infrastructure study. Join us remotely. Or not. Who hears? Beijing? Hello?
[01:35:50] Pei Jin: Hello?
[01:35:51] Session Chair: Yes. We hear
[01:35:52] Pei Jin: you. Yes. Cool. Can you see can you see my screen share?
[01:35:56] Session Chair: Nope. It reaches empty.
[01:36:00] Session Chair: So how about I share the slides and I give you control?
[01:36:05] Pei Jin: Oh, yeah. That should be pretty cool. Okay. Cool. You have control. I see. Okay. Yeah. Good afternoon, everyone. Yeah. Thanks. Thank you, Suraj, for, like, letting me to host remotely. So this is Pei Jin from University of Maryland. And today, I'm going to introduce our paper RTC Relay Server Infrastructure Study, which is a joint research project done with Meta. So okay. I see. The animation is just not here, but we're we're easy to go with. So, basically, today, we have popular RTC applications. For example, that it covers our daily life, like chatting with friends, conferencing, interviewing, video sharing, etcetera. These apps like Discord, Messenger, WhatsApp, FaceTime, Google Meet, Zoom, etcetera, they cover like, they they sync it into everyone's daily life. So, basically, from user or risk researcher's perspective, we always see the client's end. So, basically, the the apps, they deploy our client phones or iPads or our laptops. But which in between is the RTC infrastructure. It remains opaque, and it is, like, always hidden from the companies to the users. So in this paper, we're going to look into what's happening between the client when we make the call. So an interesting question is to look into is by looking into the media pass. So, basically, let's say media pass, we see how it is established. For example, we have two RTC clients. Though the animation is not here, so the picture would look a little bit messy, but we can still see that it will first try to look into whether the two clients are publicly reachable. So if it is, then PDP connection is preferred to be established to transmit the media packets. And if the the two clients, the public IP address is is not reachable from each other and for back sessions through the release of a or we say the turn server is established. So it means that it's serving as a proxy or a bridge that every media packets goes through this release server and back and forth between these two client devices. So in this paper, we have the background to see, okay, what RelayAuthorrent server is. And the next question we intuitively ask is, if you are running a RTC service or RTC company and you have a budget of, let's say, 10,000 servers in hand, how are you going to deploy them? Or we say, how are you going to build the release of infra? So generally speaking, there are two methods of doing so. So one is more centralized, which means that you
[01:39:41] Anthony: can
[01:39:41] Pei Jin: potentially have, like, longer core latency because it is potentially not that easy to get a relay near by hand. And the good thing is, for example, it is easy for loads balancing, and it is easy to maintain because there are fewer nodes physically. And the other way is we do get a more scattered solution. So for example, let's say we can distribute them, like, across 10 or 20 cities across the country or the globe. So this will give you potentially lower call latency because it would be easier to get release of, like, in the other hand. But the cons are there could be, like, higher maintenance challenges because you have more physical nodes to maintain. So these things, they are all, like, theoretical to happen. So let's see what the companies are doing. So in order to build such study, we study core problems. So first thing is the relay map. So we're we're going to plot the maps to see where the servers are located. The second thing is when we get this map, it's easy to ask. Okay. So we got a map. So which relay in this map or among all these cities or all the nodes is going to be allocated to serve this particular call? And the third thing is we have so many relay nodes in a world. So are there any plan b's if the current relay goes offline your call? So the fallback basically measures how robust the call is. And in this paper, for now, we cover three popular RTC apps, which are Discord, What's Happened, Zoom. And for Zoom, we cover both the free and business accounts to show the comparison. And we'll have more apps, like, to be covered in the future work where which will be an interesting angle to dig in. So, basically, for the measurements and challenges excuse me. Is is can can we switch the into, like, the the animation way so that we don't have the
[01:42:00] Audience Member: coverage?
[01:42:04] Session Chair: Sorry? Yeah. I don't think we can change the slides.
[01:42:11] Session Chair: I can't he has control. Yeah. Do you want me to take back control? Because you have a control.
[01:42:21] Pei Jin: Yeah. I mean, that the animation is gone, so there would be, like, slides covered with the
[01:42:26] Session Chair: With the table? What do you mean? Oh. No. We cannot change this.
[01:42:37] Petia Vasilova: It's a
[01:42:37] Session Chair: it's a it's so that's a static PDF. We cannot change the slides, and that the table is covering parts of the slides. It's I mean, we are I think we should do it on the audio track.
[01:42:48] Pei Jin: Okay. Let's let's just see, like, if we can just move go on. So basically talking about the measurement challenges in order to proceed with such study. So for one thing is we need to make enough calls from all over the world. And the solution is because we we cannot fly or, like, 10% to anywhere in the world. Right? So, basically, we use the cloud infrastructure. And the solution is we use our phone to be connected to certain Azure VMs distributed globally, which are listed in the map and the table here. And the second thing is we need to make a lot of calls. So, basically, we need to scale up the repeated number of calls. And the solution here we use is we use the Xcode app installed on the laptop to automate the tapping of the picking up and hanging up buttons on the phone's mobile front end. So, basically, the measurement test that looks like this. We have MacBooks. We have x one installed on through the caller and the caller phone. And most phones, they have WireGuard VPNs connected to the Azure VM. And on Azure VM, we have wire shacking stored so that the packet traces can be collected. Using these RTC traces of the PCA files collected, we analyze it, and we study, like, where the what the package they are sending towards, and we diagnose, like, what the release of IP is. So after we get this release of IP, we need to identify the locations. So we have three steps here. So for one thing is we study the IP intelligence services, for example, like, info.io to see what it tells us about the release server location. And after that, we need to further verify that it is true. So we carry out two steps. One is we verify the geolocations by latency triangulations using the VMs we have in hand. The the third step is we trace route to the IP to verify any city for information alongside. For example, IID means the, yeah, Duluth Airport at DC. So this information can also help us further provide further confidence of the measurement result. Okay. So okay. There could be some display issue here. Like, for one thing is we show the Zoom free account for lane map. So in our study, we find out that Zoom provisions two tiers of relay infrastructure, and the business accounts we show here has substantially more relay nodes than the free account. So we observed only two Relay locations of Zoom free account, which is located in The US East and West Coast. And both of them are run by Zoom itself. Okay. It it seems that there's, like, some display issue here. May I request you, like, share my screen again? Like, the the diff
[01:46:09] Session Chair: Can you reshare the screen?
[01:46:11] Session Chair: Yeah.
[01:46:20] Session Chair: Can you try again?
[01:46:41] Pei Jin: So so basically speaking, right, for the Zoom really map, we only have two nodes observed globally. And for the business accounts, we have substantially more nodes observed globally. We have 11, like, from our measurement scope. And note that, like, all of them, they are run by Zoom itself. For this code, the thing to highlight is it uses a hybrid relay infra. For example, 80% are run on CloudFare, 16% on Google Cloud, and 5% nodes are run on i3d.net, observe only Tokyo. For WhatsApp, it exhibits the broadest Relay footprint among all the studied applications, which have 25 Relay locations globally, and all of them are owned by Meta platforms. So from the targeting point of view, Zoom and Discord will see that they run under a fixed caller region, and they are repeated calls selecting the same IPs or IP prefixes of repeated calls. And Zoom unlike Zoom and Discord, WhatsApp, on the other hand, adopts a joint caller callee aware targeting. So it allocates three relay candidate IPs through three stem binding requests at the start of the call, and one relay is picked based on the end to end latency measurement through the stem binding request and response. So when talking about the fallback, we see that for for Zoom and basically, we count the number of sessions, the UDP sessions and TCP sessions to see how robust the call is. So counting them together, Zoom free account has one UDP session and one TCP session. Discord has only one UDP session, which makes them both non robust. And for Zoom business, it has one UDP session. And if you block that, there'll be multiple TCP sessions, including one on the original Zoom relay and multiple on AWS fallbacks. And WhatsApp has three plus three, which makes them more robust than the other two. So when after studying the locations and the targeting of the strategy of these really applications, we then see the latency match measurement, which is a result of the previous. So we see that our first thing is a denser relate deployment always brings down the core latency. For example, in this figure shown here, the darker or the deeper color of these boxes shows a higher latency. And we see that for Zoom free, the general color is the most deep deep. And for WhatsApp, it's light. So it just gives us a general latency comparison, which is Zoom free provides higher latency than Zoom business in Discord and higher than WhatsApp. And, also, we show that geographically close calls are primarily more benefited through, like, more really nice deployment. So if we see on the diagonal of these figures, we can see that, for example, like Zoom free has only two nodes globally. The diagonal is most deepened. And for WhatsApp, it is most light. So, basically, it seems that it's not always beneficial, for example, if you are going to have a shorter latency to prescribe to subscribe into a higher level of account. If you are both, for example, living near the relay, the caller, and the caller, it is substantially not necessary to purchase such higher services.
[01:50:31] Danish: Just mail.
[01:50:32] Pei Jin: So running in conclusion Yeah. No. Let Okay. We in this paper, we see that denser related deployment and smarter targeting both reduce the core end to end latency. And the hosting models, they can differ from the company to company, and the fact robustness varies. And interestingly, these results can change. So our study is measured through August 2025 to March 2026. The reported numbers may shift as the providers reconfigures for time and space. And just a lot of future work that can be done, like, after this. The first thing is we need to make repeated calls over time and space. The second thing is we can try to extend the measurements nodes to a broader scope and set and then studying different devices can also be beneficial. Last but not least thing is we're welcome to follow-up work on more platforms. Okay. Thank you. So happy to take any questions if there are any.
[01:51:36] Session Chair: Thank you very much.
[01:51:46] Danish: Danish from MPI. I have just one question regarding the using Azure as your vantage points. How that can skew your results in the at at the end?
[01:52:01] Pei Jin: So so, basically, we we say that there there is 100% the chance that there are relays that we are missing, like, in this part because we're using Azure, Only certain notes of them. So in all paper, we just bring, for example, like like a to some level of, like, comparison between the the the companies. For example, let's say, for Zoom free account, we only observe two. And for Zoom business, we observe more. So in in this matter of fact that we cannot guarantee that we opt we observed every single relay and every single IP in the world. But we just say that using the cloud infrastructure, it gives us the, like, ability or versatility to make phone calls across the globe and detect the release distributed globally. So I wonder if this answers your question.
[01:52:58] Session Chair: Yes. It
[01:53:00] Innocent: does. You
[01:53:01] Session Chair: very much. Any further questions? If not, then let's thank the speaker once again. Thank you. And now, last but not least, a short paper by Innocent. He is with University of Washington, and will talk about Towards Extensible, Auditable, and Modular Measurement specification in execution in computer networking. And we wait until the slide deck is shared.
[01:53:33] Innocent: Can everyone hear me?
[01:53:34] Session Chair: Yes. Perfect.
[01:53:35] Innocent: Is it possible for me to screen share?
[01:53:38] Session Chair: Yeah. Zoom is the share is currently enabling the share. Waiting. Now we see the slide now we see your screen, actually. Perfect. And the slide check. Okay.
[01:54:06] Innocent: Great. Alright, everyone. My name is Nissen Obi. I'm a PhD candidate at the University of Washington presenting a paper done with my, collaborator, Siak, who, was an undergrad researcher with me. Experiment measurement are very two overloaded words in computer networking, and so this is something of sort of our attempt in building an experimentation platform to add some some logic and reasoning to this work. So first, we're gonna talk about the semantics of measurement. In active measurements, say we have access to a ping resource through some programming interface, the the black box you see on the slide, by constructing probes. We can think of the white boxes, north of the interface as the client side and the white boxes south as a service side, the north and the southbound interfaces. We can present this capability as a relation. So the interface, the ABI on the slide, is a subset of the Cartesian product of a probe and a ping. And so each ABI is uniquely defined by a probe and a ping resource. But there are many such things that we can use to ping, and so this is a little bit unstructured. So let's generalize this a little bit more. In this case, all the information that we were hiding previously in the ABI is a little clearer. We have another white box called scamper that is south of the probe and north of the ping resource. And now there are two new interfaces. The ABI is now an expression expressing a relation between pro camper while the HAL is expressing a relation between scamper and PINK. Now the distinction between the client side and the server side is a little clearer. One thing to note in this formalization is that we can express the relationship, by the Cartesian product of the white boxes or by the product of the relations, the black boxes, these these interfaces. And the key idea is that the relational point of view is a more powerful abstraction. You'll see kind of where we're going with this. So say you don't want to use Scamper, but you want to use something else that supports pings. And so we want to be parametric over the implementation of the resource like ping. This motivates the addition of a resource white box, and new north and southbound interfaces around that. And so this generalization abstraction allows us to discriminate and decompose context. But as we're doing that, we're adding fixed points in terms of our our abstraction. And so as we generalize and we create these white boxes and interfaces, the interfaces act as guards for our white boxes, guarding the north and the southbound interfaces, playing the role that a protection ring would play in, let's say, an operating system. And so, for example, if we wanted to switch out, let's say, scamper for the default Linux ping utility, it would be as simple as just in that component. And we can see that the impact of this change is local, and can be done without breaking policies defined at higher levels or commuting with those defined at lower levels. So the interfaces become places of defining behaviors between these white boxes. So MTL, which is the work that we do and we present in the short paper, is sort of a universal adapter. And so we'll walk through this kind of diagrammatic calculus, you know, demonstrating what it would mean to specialize an MTL ping resource. And so our base MTL template is free, so it has open incoming wires for probes and outgoing wires for resources, meaning that it captures the entire universe of probes and resources. We specialize this by defining a resource interface called our ABI. You can think of this as ping. And now, like in algebra, we can join and compose these two together and match and identify the resource wires with the ping interface wires. And so the result is an MTL module template that is specialized for a Ping resource. We have reduced the universe of possible resources to those that implement Ping. On the other hand, we have the same Ping interface, but now we have Scamper running on our Linux machine, for example. And so Scamper has its own API that allows it to be interacted with. And to make this empty already, we compose Scamper with our ping interface. So the result is an interface constrained Scamper. Now since our previous resource specialization has a southbound that is a ping interface, and our now constrained scamper has a northbound, that is a ping interface, by definition, we can combine these common north south north south interface. The result is a scamper resource module, which fully implements ping that can receive ping valued probes. That is a specialization of the southbound works inversely to restrict the universe of possible probes that can be executed on this resource. And so in our short paper, we kind of motivate
[01:58:40] Danish: Oh, sorry, Innocent. You're way out of time now.
[01:58:46] Innocent: Apologies. I thought I had
[01:58:48] Session Chair: You are eating into your question time. If you wanna go ahead and summarize, like, you can go ahead and close. Yeah.
[01:58:53] Innocent: And so the way we do this in the paper is we motivate an essential, like, and, two, two semantics for for the for the context. This looks something like this. Basically, we can turn our language semantics into these diagrammatic calculi we use. Our motivation is to think of ping as operations, versus the usual way of thinking of ping as as utility. And so our overall view is that there's a deep algebra underlying this that thinking of ping as operations motivates thinking of these algebraic signatures and Ping as utility as these algebras that implement, our our signatures. And so we can think of, our measurements and our measurements, experiments or measurement scripts as these implementations. So I'll stop there, and, yeah, we'll open for questions.
[01:59:42] Session Chair: Thank you very much. So any questions? Maybe then just a simple question. I mean, that sounds all very theoretically. Is there any concrete implementation? I mean, see some code here. It's a source implementation that can be downloaded.
[02:00:08] Innocent: Yeah. So we don't have an we don't have a public implementation yet, and so that's what kind of we're working towards in our paper. But we we have a proven model that's worked in our in our lab environment. And so what we're doing now is, working on releasing a Go library that has some of these semantics built in. But we have them demonstrated kind of lifting a scamper, implementation and running pings on scamper using this framework.
[02:00:33] Session Chair: Okay. Cool. Looking forward to the extension. Then if there are no further questions, let's conclude the longest session of this workshop. Thank you all. Enjoy the break and come back half past four to the session on Transport and Media.
[02:00:50] Session Chair: Thank you, Matthias. Thank you very much. So in other conferences, when
[02:01:05] Session Chair: you have no
Session Date/Time: 20 Jul 2026 14:30
[00:00:07] Ali Begen: Oh, you you you
[00:00:08] Thomas Witt: opted for the healthy stuff.
[00:00:10] Suresh Krishnan: I didn't have time for anything
[00:00:12] Ali Begen: else. Like,
[00:00:12] Sebastian: there's too much
[00:00:13] Suresh Krishnan: of a line for everything.
[00:00:16] Matthijs Engelbart: Not
[00:00:19] Presenter: not good.
[00:00:21] Thomas Witt: So, I mean, if if Corey is not appearing, which I don't know, then then I'll just do it.
[00:00:27] Suresh Krishnan: Yeah. Go for it. Yeah.
[00:00:31] Thomas Witt: It's time. Right?
[00:00:32] Dirk Kutscher: Yeah. So let me start.
[00:00:45] Thomas Witt: So welcome to our last session. Please join in. Find your seat.
[00:01:28] Suresh Krishnan: Can you pick up the clicker? You can pick up the clicker. Yeah.
[00:01:35] Thomas Witt: So I welcome you to the last session on transport on media and to start with a real surprise for those who didn't invest in SpaceX, we are going to space with Quick now. So welcome Natalia for her talk.
[00:01:56] Natalia Petroski: Thank you for the introduction. Hi everyone, I'm Natalia Petroski. I'm also from the Technical University of Munich, and I'm here to present our work on QUIC in space initial measurements and analysis. So reliable communication is, of course, always critical, but it's especially critical for space missions. There, however, there are some challenges that we face when we are communicating in space that we do not encounter on terrestrial networking. Amongst those, not extend not not extensively, for example, intermittent connectivity, and high latency due to orbit mechanics and low data rate and asymmetrical links due to technical limitations. The current approach to deal with these challenges is the bundle approach, which is pioneer being pioneered by the DTN working group here, the ITF. However, in recent years, and sort of older or another approach has resurfaced, and that is to ask the question whether we can deploy the IP stack in space and with it quick. This this topic is being dealt with by the TIP Top Working Group at the ITF and is also what I'm dealing with in my research. Specifically, I'm trying to answer the question, how well do QUIC implementations perform in space? So Okay. So QUIC itself is the focus of the TIP Top Working Group, and this is also why we are focusing on QUIC, and it has already proven its viability in initial simulations. We want to look at QUIC as well as two protocol extensions, how well they might affect communication in space. Notably, the frequency protocol extension and the careful resume protocol extension. Frequency allows us to delay acknowledgments, which should be, at
[00:03:43] Matthijs Engelbart: least in theory, beneficial on highly asynchronous links, which is the case for space communications, for space links.
[00:03:50] Natalia Petroski: Other work which I'm, yeah, referencing here has already proven that this protocol extension is useful for gear connections. And in order to delay these acknowledgments, we can play around or vary two variables, notably the max ack delay and the ack threshold. So the max ack delay denotes how long we have to wait between or we are allowed to wait between, yeah, two acknowledgments, and the argThreshold tells us how how many ack-eliciting packets can be received before an ack is issued. The other protocol extension we want to look at is the careful resume extension. This protocol extension allows us to cache congestion control parameters, notably the congestion window and the minRTT, in order to shorten the slow start phase occur for a recurring connection, which also, in theory, should be beneficial for space links. In order to meaningfully evaluate this, we will do a statistical evaluation, which I will be explaining in just a second. So for measurements, as you probably are aware, there's a lot of QUIC implementations out there we had to focus on too. We wanted to focus on too, notably Quinn and Quiche. You can see the concrete version numbers on the slide here, both are written in Rust. And we build our own workbench, which is based on NetEm and Linux namespaces and looks as you can see here on the slide. So we have a client and a server communicating over a singular bottleneck link. To this bottleneck link, certain parameters are applied, which you can see written here as well. So for the data rate, we assume 25 megabit per second. The rationale behind this value is that this is the value of the intersatellite link of the Iridium satellites.
[00:05:32] Matthijs Engelbart: And
[00:05:32] Natalia Petroski: then we play around with the delay values, which we set to 0.5, 1.5, and three seconds, which translates to an RTT of one, three, and six seconds. The rationale behind these values is that we want to focus on communication to the moon and slightly beyond, as this is also the focus of the TIPTOP working group. And for the limit, we set the BDP of the link, where we also calculate any outage durations, where we take it into the calculation here as well, which, of course, is only possible since we know in advance how long any outages are going to be, which is obviously not quite too realistic. So I already mentioned that we want to look at the protocol extensions, but another topic that we also want to focus on in our evaluations are the outages due to the intermittent connectivity. For this, we have to introduce two new metrics. First, we have to introduce the stability metric in order to determine when a connection would reach a sort of stable state after the slow start in the beginning of the connection. The reason we wanted to do this is we wanted to have outages occurring in the stable state. And, well, to know when this happened, we had to introduce this metric. So I'm not going to go too much into detail how exactly this is being done because that would need a longer explanation, which you can see in the paper. But we assume a connection to be stable once the throughput stays relatively consistent over a period of time. And then in order to understand whether or not a connection has recovered from an outage, we also need a recovery metric. For this, we compare the congestion window just before and after the outage in a jumping window of 10 times RTT. So we calculate the average congestion window in this time period, which we then refer to as the target congestion window, and then calculate it after the outage as well, and once the congestion window in these jumping window sizes has reached the target congestion window, we understand the connection to have recovered. We translate this, or we understand this as the both the recovery percentage, which is simply the fraction of this recovered congestion window and entire congestion window, as well as the recovery timestamp, which is simply how long it took to reach, yeah, any a recovered state after an outage. And it's always relative to the end of the last outage. In case no recovery was observed, which we're going to see in just a second, we denote this value to be minus one. In order to have some sort of confidence in our results, we evaluated this over three different runs and calculated both the mean and the standard deviation values of both the recovery percentage and the recovery timestamp, of which I will only presenting the mean values here in this presentation. So moving on to to the results. First, we played around with the ACK frequency protocol extension. This is also the only, scenario where we actually use asymmetrical links, notably with a factor of 10, meaning that the downlink had a data rate one tenth of the uplink. However, even though we played around with both the max ACK delay and the requested maximum ACK delay and ACK threshold, sorry we could not find any significant improvements in the flow completion times when using this protocol extension, which is also the reason why we only tested this in Quinn where it was already implemented and did not implement or evaluate it in Quiche. Moving on to the careful resume protocol extension. Here, I already mentioned earlier that we conducted a statistical evaluation. This means that we had 10 baseline runs which did not use the protocol extension, but simply cached the values, and then 10 careful resume runs which use the protocol extension with the parameters that we had just cached in the baseline runs. And here, you can see the results on the slide here as well. So we could see some improvements for the flow completion time for Quiche, however, not really for Quinn. For Quinn, it was especially surprising that BBR that the flow completion time of BBR actually worsened when using the careful resume protocol extension. However, we chalked it up to an implementation error, since in previous or, yeah, sort of very initial evaluations, this was not the case. And so we chalked this up in this case to simply an implementation error, which we will have to fix. Moving on to our outages here, we evaluated both singular and multiple outages, which you will see on the next slide, and, yeah, represented it in heat map heat maps. So here you can see, in each field, you can see the in the color, the recovery percentage, which is always the average recovery percentage over the three different runs, and, the timestamp, which is the value written in each field, again, the average value. So here we can see, that the recovery percentage was more stable for Cubic than for Quinn as the colors for the recovery percentage stay closer together for cubic than they do for Quinn, for BBR. The recovery time stamp, rose for higher RTT, which is as expected. However, what is unexpected is that there are some exceptions for an RTT of three seconds, notably for quiche cubic where this is persistent or consistent. Looking more closer at the values, we can see that the only sort of combination of implementation and congestion control algorithm that struggle to recover is that of quin cubic. As you can see that the this is the only combination where we can see minus ones in the values for the time stamps, notably also only for lower RTTs, which could be kind of which is kind of surprising. Another sort of oddity we can see for, Quinn is that for Quinn BBR, we see a very high spike, for an RTT of one second and an outage duration of eight seconds, so that yellow field that you can see down there, which is simply due to how works. And this is this is understandable once you look at the exact sending behavior of QIN that it sends in certain bursts. Moving on, before I run out of time, we also evaluated two outages. So these were, again, both in the stable state mainly. So the first outage here occurs at the same time as in the previous evaluation, and the second outage occurs once the recovered test is understood to be recovered. The exact values for this can be found in the paper. So here, we can see that the recovery percentage now stay more stable for KISH rather than for q for Quinn as, again, the colors are more closer to each other. And the recovery timestamp, again, grows for high RTT, which, again, is ex as expected. But, again, there are some exceptions for an RTT of three seconds, which in behavior is quite comparable to that of a single single outage. And similar as before, cubic targets with recovery, notably now, not only for Quinn, but also for Quiche, but for Quinn, we can now see that that we do not achieve recovery for higher RTTs, but we struggle to achieve recovery for lower RTTs. This already brings me to the end of my presentation. So we have looked at different protocol extensions and how outages might affect quickcom quick communications. We saw that the ACK frequency, protocol extension, unfortunately, does not seem to be beneficial for deep space links, at least not in our evaluations. The careful product the careful resume protocol extension shows promise, even though there were some oddities that we need to figure out, but we chalked it up to implementation details. Regarding outages, we could see that cubic struggles with outages, especially in QIN, which is quite expected. And we could also see that the recovery time stamps are relatively stable regardless if one or two outages occur, especially if these outages are both in a sort of, yeah, understood stable state. Thank you very much for your attention.
[00:13:36] Gorry Fairhurst: So, Gorry Fairhurst, and thanks ever so much for bringing this. I suspect that the bugs and implementation things are a feature of actually testing like this because when we did implementation in Cloudflare Quiche of careful resume, we also saw timing errors that we didn't expect, which we traced down to unusual ways in which times were manipulated and also recovery problems. So this is actually stress testing protocols, so it's a good thing to do experimentation with, and I really think it's good. Did you ever look at happy eyeballs?
[00:14:11] Natalia Petroski: Go ahead.
[00:14:12] Gorry Fairhurst: Happy eyeballs and the idea that QUIC and TCP might race one another over this sort of path.
[00:14:19] Natalia Petroski: Sorry. Happy eyeballs.
[00:14:21] Suresh Krishnan: Happy eyeballs.
[00:14:22] Natalia Petroski: What is that? It's like Maybe we'll talk about this.
[00:14:26] Gorry Fairhurst: Yeah. The answer is probably no, but that also might be an interesting thing to stress test. So thank you.
[00:14:36] Robert: Hello. My name is Robert from Hasselbladnik Institute. On slide four, you had a little bit of assumptions on your bottleneck parameters in starting. Could you elaborate how you got to those?
[00:14:49] Natalia Petroski: Could you repeat that, please? I didn't
[00:14:51] Robert: catch You had assumptions on your bottleneck parameters. Uh-huh. How did you come to those?
[00:14:57] Natalia Petroski: Oh, I mean, for like, what specifically? For the the data rate, that was, again Yes.
[00:15:03] Matthijs Engelbart: For the data rate.
[00:15:04] Natalia Petroski: The data rate, yeah, that's that's simply the intersatellite link of the data rate of the intersatellite links of the Iridium satellites. So this was our next best guess. It was very hard to find good values to use there, and this seemed like a sensible value. So that's why we used it here.
[00:15:19] Robert: Alright. Thank you.
[00:15:23] Danish: Hi. Danish from MPI. Regarding the BBR experiments that you've done, you know that BBR has probing phases like RTT probing or bandwidth probing. How your outages were were, covering these probing phases?
[00:15:40] Natalia Petroski: They didn't and they didn't take that into consideration on what exact phase the the, yeah, the PBR currently was. So that wasn't taken into consideration. Thank
[00:15:52] Suresh Krishnan: you.
[00:15:57] Thomas Witt: You want to come up? Or
[00:15:59] Suresh Krishnan: Yeah. Okay. Buddy, you wanna run the slides?
[00:16:10] Natalia Petroski: We wanted to take shot. Yeah.
[00:16:17] Daniel: Everyone. I'm Daniel. I'm a PhD student and research assistant at the Technical University of Munich. And today, I wanna talk about how h three prioritization can enable the cooperative scheduling in MASQUE. So before we start, though, let's assume that we have a client and server here that are communicating over some kind of legacy protocol that is difficult to extend. And even though we may not be able to extend this protocol directly, what we can do instead is that we introduce a MASQUE proxy through which we then tunnel all of our traffic. And this works because when we operate MASQUE over H3, we end up inheriting a series of capabilities and extensions that are provided by the entire underlying stack. For instance, we may have selective reliability. We can have prioritization mechanisms and even support for multiple transport layer paths at the same time. And what these features here have in common is that they all ultimately rely on some kind of tailored scheduling solution that relies on application specific knowledge that has been re that has been provided by the client, and yet MASQUE at the moment does not provide a standard way to do so. The case that we wanna make today is that the extensible prioritization scheme for HTTP or EPS for short is already in fact capable of providing these missing semantics. UPS, operates on two parameters. We have u and I, where the u stands for the urgency, which is a priority that ranges from low to high. And we also have a Boolean flag, which which essentially, requests the peer to multiplex its resources. Beyond that, we can also introduce new parameters that are able to extend these that are able to extend these scheduling capabilities as implied by the name, all of which in a mass context and ultimately end up applying to data streams and also connect requests. When we looked at related work, we found that both cooperative approaches as well as flow-aware scheduling were able to provide performance enhancing services in MASQUE. What was missing, though, was precisely this signaling protocol that enables the clients to convey its intent to the proxy, which is why in our work here, we ended up repurposing EPS to that end. So, EPS, it was very natural to embed these cooperative semantics into EPS because it essentially already has these strict priorities that we see in Multipath QUIC literature. These are covered here by the urgency parameter, the demand that we have for multiplexing in QUIC, and by extension, H3 and MASQUE. These are provided here by the incremental flows. Nothing in the specification prevents EPS from applying to multipath topologies either. And by keeping these new extensions that we introduced constrained to what we actually control, so namely the client and the proxy here, we don't have to involve the server at all, which is which was a limitation in prior work that we had seen with EPS extensions before. To actually then go ahead and evaluate this approach, we designed and implemented our own MASQUE implementation, which we called MASQUE-NADA. And this it is essentially able to provide a scheduling framework that enables you to evaluate and measure these dynamic approaches in a reproducible manner. To do so, we built it on top of tokio-quiche, which is a quick library that made it very easy to essentially implement these reliable and unreliable MASQUE flows that we have in MASQUE as well as shipping with this EPS signaling out of the box. Beyond that, we also wanted to experiment with Multipath QUIC, and there we found a draft implementation for the protocol extension already available in the pull request. Our proxy is also able of to offer a pluggable scheduling interface, which differs from the baseline that we previously had, where that proxy was essentially working with a fixed and static scheduler for all of its connections regardless of the the preferences that the client actually had. So our contribution was to essentially move these scheduling decisions to the connection level and beyond that, also to enable the client to request a custom scheduler that is being offered at the proxy using this new EPS extension there, which essentially then operates on top of this outer Multipath QUIC connection. And we also looked through several QUIC schedulers from the literature, some of which we were then able to essentially modify to accommodate these EPS semantics, and that resulted in schedulers that we here call EPS aware. We then evaluated this whole idea on top, of a proper bare metal testbed. So there we had a client, a proxy, and a server. We had two physical links between the client and proxy, which then maps to these logical paths at the transport layer. And the traffic that was flowing was initiated through the client, which essentially established nine flows over that proxy. And the client also set these EPS requests, EPS parameters that essentially were understood here by the proxy, which in turn divided these end to end flows that each retrieved a one megabyte file into these groups of low, medium, and high priority. And we also set the incremental flag as well as varying these scenarios with respect to whether the tunnels were either reliable or unreliable MASQUEs. So we were essentially varying the substrate. We also evaluated different path conditions. So we either had, identical path properties or heterogeneous conditions where one path was significantly slower. And we, of course, also evaluated different scheduling approaches, including NDPS Aware scheduler as a case study. What our results then showed at a high level is that for one, with unreliable MASQUE, these u and I semantics, they are effectively lost. So the flows, they were being essentially scheduled in a manner that was evenly splitting the throughput, which is in line with the specification, though, since according to the specification, they should only apply to streams. And that brings us to reliable MASQUE where we have a stream as a substrate, And there, the urgency parameter was indeed able to lower the completion time of these prior prioritized flows. So the higher that we set the urgency, the faster these flows actually completed. And the incremental flag there that we had set, it was essentially making sure that then we didn't see any blocking issues for these flows that were being sent in a parallel fashion at the same level of urgency. On top of that, our proxy was indeed able to parse this new EPS extension that we introduced, which at the end of the day is what made that on the fly scheduler selection possible, which is why we see here that even for these unreliable MASQUE flows where these semantics previously did not apply, there, we're all we are able to, yeah, select the scheduler, and we ultimately saw that the round robin strategy, for instance, performed pretty well under these identical path conditions. The round robin scheduler, however, then performed much, much worse, when we actually moved to heterogeneous paths, and that's because now one of the paths was significantly degraded in terms of the RTT and its rate, which, just goes to show how important it is that we're actually able to dynamically switch away from an underperforming scheduler, perhaps due to not fluctuating network conditions, for example. Beyond that, the EPS aware scheduler that we had adapted before, we were able to see that it was indeed able to provide new functionality, addressing a recommendation in the EPS specification that these methods to address starvation should be implementation specific. And here concretely was what this scheduler was doing is that it was effectively able to sometimes yield the sending opportunity of a high priority flow so that these lower priority flows could still make forward progress while still maintaining the overall completion order that the client had requested. So this is essentially the behavior of that incremental flag, except that we are now able to sometimes skip between these urgency classes. All of this, however, is still very much a prototype, so we do want to move our implementation to a more mature, quick implementation. We want to evaluate this entire setup under more realistic traffic conditions as well next, as well as not only working with this concrete case study, but also testing other scheduling approaches, that we can make, EPS aware. One approach that seems promising that we had seen in the literature, for example, is that we, take a high priority flow and that we always assign it to the fastest path, which should reduce the occurrences of out of order arrivals, for example. Beyond that, at the moment, this extension is only, yeah, doing its job when you actually create a new tunnel. And a more idiomatic approach to this to stay a bit more in line with the EPS specification would be to leverage the priority update frames that are able to reprioritize flows in according to the EPS specification. And with that, we already come to the end. So we were essentially able to see that EPS is indeed able to provide per flow quality of service in MASQUE. EPS was also able to enable these cooperative approaches, and it does so while still allowing for room to specialize. Taken together, we believe that EPS is a suitable signaling scheme for MASQUE scheduling that we would like to see adopted more. Thank you.
[00:27:12] Gorry Fairhurst: Any questions?
[00:27:18] Lucas Pardue: Hi. Lucas Pardue, one of the coauthors on the priorities draft and a mask enthusiast. So thanks for this work. It's it's real. Like, a MASQUE proxy does have these scheduling issues to deal with, and it's especially with the multipath stuff, it's really interesting. So for the streams, it's it's kind of given. If you're in HTTP/3 and you implement prioritization at all as a server, the thing to remember is kind of the signaling is distinct from the server doing the scheduling. And the header, so the the priority header is like one of the inputs into the scheduling decisions the server can make. And in the draft itself, there's a whole load of text that is trying to guide people towards doing something that's sane for web, but that kind of all these caveats and escape mechanisms for doing other stuff. But it was written before we had quick datagrams and a kind of MASQUE flow of all of these other multiplex datagrams even though quick doesn't model it as a transport service as multiplexed flows. So it's it's very complicated and there's lots of opportunity to do better. So really, thanks for doing this research. I posted a link in the chat about a a way to articulate just Datagram priorities to try and give a semblance of this different flow linked to those connect requests. It's it's, like, similar, but slightly different than this. Might just be worthwhile looking into that to see kind of any any overlaps with those things. But, yeah, just just good work. Thanks.
[00:28:49] Daniel: Thank you. Yeah. That was a lot to take in, so I would like to connect offline before Yeah. Later on. But we we we did have a bachelor thesis that I would share that attempted to do some prioritization on top of datagrams as well. So perhaps I could look into that for more concrete details on what to do next. But, yeah, I do admit that this is all a little bit hacky and that we're, yeah, trying to push the boundaries of what UPS is supposed to do. But, yeah, I'm open to to feedback on this.
[00:29:22] Lucas Pardue: But my my only comment there is this is this like, this whole area is open for improvement, so there's no clear answer. This is why the research is so important to do. Thank you.
[00:29:31] Presenter: Awesome. Thank you.
[00:29:33] Ingrid Tovin: Yeah. Thank you ever so much.
[00:29:56] Matthijs Engelbart: Hello. My name is Matthijs Engelbart. I am also a PhD student at the Technical University of Munich, and I'm going to present an architecture for multiplexing data channels and media in WebRTC using QUIC. And this is work that I've done together with Fabian Fumboldt, who's also a student at TUM, and my supervisor, Jörg Ott. So in the last couple of years, Jorg and I have also been working on RTP over QUIC, which is just RTP over QUIC and no data channels. And that draft is adopted in the AVTCORE working group. It's been in the works for a couple of years now, and it basically defines an application layer protocol that is mostly an encapsulation format for RTP packets and QUIC. It supports QUIC streams and QUIC datagrams for different use cases, and it has a so called flow identifier for multiplexing different RTP streams within a QUIC connection or multiplexing RTP and RTCP in a single QUIC connection. And the draft also contains some applicability considerations, for example, for congestion control or for generally sending RTCP over QUIC. And an interesting use case for RTP over QUIC is, of course, WebRTC, which is a framework for peer to peer real time media communication on the web. WebRTC contains media communication using RTP over UDP, usually UDP with a fallback to TCP if UDP is not available. And WebRTC also contains data communication, which allows applications to exchange arbitrary data using a protocol called data channels, which works on top of SCTP and also runs over UDP. Something that easily happens when applications use both of these at the same time is that because the different protocols use different congestion controllers, these compete with each other for the same bandwidth in the network. So if we have an application sending, for example, a file transmission over a data channel, and we also have a media stream running in parallel, something that can easily happen is what we see here is that the media stream receives very little bandwidth while the data channel consumes a lot of the bandwidth available in that network. And that happens because the SCTP uses usually a loss based congestion controller, which also leads to the RTP stream observing very high jitter, as we can see in the bottom graph here. And a solution that has been proposed for this in the past is called coupled congestion control, which basically couples the congestion controllers of RTP and SCTP to make sure that both of these can get a fair share of the available bandwidth. And I would like to talk about a different alternative architecture using RTP over QUIC, which would use RTP over QUIC and then multiplex data channels on the same QUIC connection so that they both use the same congestion controller and share the bandwidth fairly between each other. So the traditional architecture in WebRTC looks like this on the left side, where we have IP and UDP in the bottom, and then we have DTLS on top to secure the RTP streams and SCTP, and on top of SCTP, we have a data channel protocol. And the architecture that we have been working on is this, where we have QUIC as the transport protocol, and on top of that we have RTP over QUIC, and then the same, or basically the same data channels protocol, but on top of QUIC instead of SCTP. We can reuse the RTP of a QUIC encapsulation format that we have already defined in the draft I mentioned before. We need a new data channels protocol because we cannot directly transfer the one that we have from SCTP because we need to adapt it too quick. We have done that in this individual Internet draft that I've linked up there. And then we need a actual multiplexing mechanism for multiplexing these two protocols in one single quick connection, which basically means we write a new or define a new application layer protocol on top of QUIC that contains these two protocols and defines how to multiplex them. And one idea for multiplexing them could be to use flow identifiers I mentioned before for RTP by using the same flow identifiers for data channels and then defining different IDs for RTP or data channels, or an alternative could be to use different quick stream types, like using bidirectional streams for data channels and unidirectional streams for RTP only. We, of course, also need to define congestion control and bandwidth estimation, and in real time communication, we usually use congestion controllers like GCC, SCREEM, or ADA, which are based on one way delays, which can be implemented in QUIC using the packet receive timestamp extension, which is currently work in progress in the QUIC working group. And the estimated bandwidth of that congestion controller would then be used to allocate different shares to the protocols that we have on top of QUIC. And we would use a pacer at the QUIC layer to limit the overall transmission of the connection so that none of the protocols can exceed that. We also need to talk about scheduling and prioritization. Media encoders usually are producing data at a specific rate and can be configured to produce data at a target rate. So we would have the allocated rate from the bandwidth estimation, and then we would configure our media encoder to produce media at that rate. Data channels are not necessarily rate based, so we could have a file transmission, for example, which is just a file and it doesn't produce data at any rate. It just is a file of a certain size. Would like to prioritize, or an idea that we had to prioritize these or schedule these, is to always prioritize quick streams and datagrams that carry media, because that allows us to ensure that we send media always when it is available at low latency, which is important for real time media streams. But since the media is rate limited and we only allocated a certain share to that media, it will not consume the whole bandwidth of the connection and the remaining bandwidth that is available would automatically be consumed by the data channels. And that allows us to have data channels that are elastic and reacting to what we are sending on the media side because media encoders sometimes overshoot or undershoot the target rate that we configure. For example, if you are creating a keyframe for, or if a media encoder creates a keyframe, it may overshoot the allocated target rate, and in that case, we would like to automatically decrease the data channel rate. And that happens automatically if we always prioritize quick streams or datagrams that are carrying media. And then the other way around, we would also like to use bandwidth that is not used by media. For example, if there is not enough data to be sent from a media stream, we would like to use that free bandwidth for the data channels. And if there is no media to send, by just prioritizing media, and there is no media to send, then the data channels would automatically consume that. So, we made a prototype implementation of this architecture. We are basically building an application using WebRTC, which sends video from one peer to the other, and on the side, or in parallel, it sends a data channel, either a large random data transmission or a file transmission in our experiments. We used the aiortc implementation to test this prototype, and then we implemented the alternative transport that I using the architecture that I just presented, which is implemented in a fork of QuickGo, which contains the extensions that I mentioned before for congestion control and scheduling. The both transport versions that we implemented use the same bandwidth estimation algorithm, which in this case is taken from the aiortc implementation, which is based on Google congestion control. And we evaluated the architecture using network emulation, using Linux virtual network interfaces, and TC NetEm. So here's some initial results for our prototype. In this case, it's a file transfer that starts after about thirty five seconds. So we have, in the beginning, only the media stream, which slowly ramps up the media rate at which it is sending a video stream. And then after about thirty five seconds, we start a data channel which transfers a file of a fixed size. And on the left side, we see the bandwidth and the latency that is taken by the WebRTC implementation, where we can see that the SCTP stream, similar to the graph that I showed in the beginning, ramps up the congestion window and takes up almost all of the available bandwidth. And then after the file transmission is done, after about sixty seconds, it drops because the data channel is closed. And then afterwards, the media stream can ramp up again and find the bandwidth limit to use it all for the media. On the right side, we see the same using the architectures that I just presented using QUIC. Again, after about thirty five seconds, we start to transfer the file using a data channel over QUIC this time. In this case, the overall transmission of the file takes longer because the data channel cannot use the whole bandwidth that it used before in the WebRTC implementation. But that's intended in this case because we only allocate 50% of the bandwidth that the congestion controller estimated to the data channel, and the rest of it can be continued to be used by the media stream. And applications are, of course, free to choose different allocations there. 50%, in this case, just an example. We also see some issues with the implementation that we have, but that we think is not related to the architecture itself, but rather to our implementation, because we see in the latency graph on the right bottom that there are some latency spikes, which cannot be seen in the WebRTC version. But we think these are related to buffering in the scheduling mechanism that we have, and we are currently continuing to investigate this. But we looked at the RTT of the quick connection, which actually does not have these latency spikes, and so we believe they are only inside the sender site implementation because the latency plot that we have here is the RTP latency from before passing it to the quick sender until after it was received by the quick receiver. We also ran some more experiments with different bandwidth configurations and one way delay latencies. And this one is the example for five megabits per second using different latencies. And the important part here is the orange cross, which is on the bottom left mostly, which is the RTP and data channels implementation over WebRTC, which is mostly staffed by the SCTP next to it. And the red one is the quick version that we implemented, which has a higher bandwidth, closer to the about 50% allocation. For more details of these results, look at the paper. They are also for different bandwidth configurations. Yeah. So in conclusion, we implemented the multiplexing version of RTP and data channels in QUIC, and showed that there that we can build this protocol. We implemented real time congestion controllers, in this case something similar to GCC, that can be integrated and quick using the packet receive timestamp extension. And we can see that this multiplexing can solve unfairness issues that we've seen in the original implementation at the beginning. And, yeah, that's all. Thank you.
[00:42:14] Gorry Fairhurst: Take questions. I've got a question to start with. Did you look at partial reliability, ever, when you were thinking about this, because QUIC offers different types of service?
[00:42:24] Matthijs Engelbart: Yeah, we were thinking about it. We didn't test it exactly in this one. We were using QUIC streams here, so there are retransmissions for the RTP packets. We are canceling these streams after a timeout so that the retransmissions does not continue to happen forever. If the frame is too late, it is just dropped. But we did not look at the partial reliability in detail here.
[00:42:49] Gorry Fairhurst: Thank you.
[00:42:53] Francois Michel: Hi. I'm Francois Michel, Apple. Thanks for the presentation. Super interesting. And I was wondering how applicable it would be to web transport in addition to QUIC, if you thought of that, and if there would be any limitation to the approach over web transport.
[00:43:12] Matthijs Engelbart: So I didn't get the last part of the question, but for web transport, I think there were also discussions about web transport when we were discussing RTP over QUIC and AVTCORE. And currently, RTP over QUIC is only defined for QUIC and not for web transport. I think that work would have to be done. It might be applicable to that too. I I don't really know how WebRTC WebRTC would fit into WebRTC in general, but could could be discussed again, I guess, how this one applies to web transport.
[00:43:42] Francois Michel: Okay. Thanks. Yeah. The the what comes to my mind is that if you want to run this in the browser, it would be easy with WebTransport. You have similar semantics. But yeah. Thank you very much.
[00:43:54] Matthijs Engelbart: Yep. Thank you.
[00:43:56] Sebastian: Hello.
[00:43:59] Dirk Kutscher: Question on slide number nine where you had the comparison. It wasn't obvious for me to see which is actually better since the ranges on the y axis are also different. You have like
[00:44:16] Matthijs Engelbart: You you mean the latency? Latency. Latency. Yeah. Right. That's what I meant. It's so I I would say the spikes on the right side are definitely not better. If we look here, we can see that for different configurations, the latency varies less for the quick version. That's what I meant on this slide. The spikes are, of course, not better, and I think that's an implementation issue that we have. But it needs to be investigated further. Yeah.
[00:44:52] Sebastian: Do we have time?
[00:44:56] Dirk Kutscher: So I did I get this right that you did not use quick streams for for the encapsulation?
[00:45:02] Suresh Krishnan: Or
[00:45:02] Matthijs Engelbart: We we did use quick streams.
[00:45:04] Dirk Kutscher: Okay. Okay. Well, thank you.
[00:45:09] Gorry Fairhurst: Thank you.
[00:45:36] Sebastian: Sebastian?
[00:45:44] Sebastian: Expert spot.
[00:45:45] Suresh Krishnan: Yep. You can move around. I think the camera tracks you. Don't go too far. Not that way.
[00:45:51] Sebastian: No. Okay.
[00:45:52] Lucas Pardue: I'll try.
[00:45:52] Suresh Krishnan: And don't fall off.
[00:45:53] Sebastian: I think it's important. I I would try. So good day good afternoon, everyone. My name is Sebastian, and I'm PhD student at the Gdańsk University of Technology this time, not Munich. And today, I wanna talk to you about much better. Today, I wanna talk
[00:46:11] Sebastian: to you about
[00:46:11] Sebastian: Low-Latency Live Streaming over
[00:46:17] Dirk Kutscher: MoQ.
[00:46:18] Sebastian: Why doesn't it work? Oh, there we go. Okay. So in a live streaming scenario, we have a network viewer who wants to see the live edge. So what is happening right now? And we want to keep that viewer as close as possible to that live edge, so we have only very little buffer. However, when the network cannot keep up with the data, we eventually end up accumulating backlog, so that the viewer falls behind. And if
[00:46:52] Suresh Krishnan: okay.
[00:46:52] Sebastian: Okay. And if the viewer eventually wants to return to the live edge, we have to drop that backlog somehow because video is produced at the same rate as it is consumed. So what you actually drop and how you drop it decides what your what the viewer sees, and this is what we are discussing today. So one protocol which has been designed with that scenario in mind is Media over QUIC. So let me just introduce what Media over QUIC is. The goal is to have CDN style scalability compared with low latency. It uses QUIC, as the name says, to leverage multimedia, and it's a publish subscribe protocol, meaning you have a publisher who offers media, you have a subscriber who wants the media, and you have Relay which can which can forward it and fans out to achieve CDN size scalability. The subscriber who wants the media, it subscribes to an identifier called a track. A track in the end is nothing more than a sequence of groups. Each group can can can consist out of one or more subgroups. Each subgroup gets mapped to one quick stream, so you have some concurrency in there. And each subgroup is then a sequence of objects. To keep the view of the subscriber close to the live edge, we can have delivery time outs in per subscription, which are measured on the object and executed on the subgroup. Meaning, each object gets a certain time budget, and if you cannot forward that object within said time budget, you kill the whole subgroup that that belongs to the object. So Media over QUIC works really well for video streaming because it mirrors the the structure that video already has. Because the video is, in the end, nothing more than than a sequence of so called group of pictures or short GOPs. Each GOP is now a sequence of frames. The first frame in the GOP being a keyframe being independently decodable. And then the p frames are then depend on either the keyframe or other p frames. So this mirroring allows us for a really, really easy mapping between the two, where each video becomes a track, each group of pictures becomes a group, and each frame becomes an object. And in our setup, we don't use any internal parallelism where we could use subgroups, so we are only using one subgroup. And for the rest of the talk, they will be synonymous to me. So imagine now that you are a quick stream, and you are working on delivering this or sending out that keyframe to our to the viewer. However, something something happens. It takes so long, and the media of a quick application now says, okay. Time's up. You have to reset the stream. What that means in the scenario, you end up with nothing at the re at the viewer because you have canceled stream, and you'll only have a partial keyframe, meaning you have nothing. So congratulations congratulations, you've wasted resources, and you've gained nothing. So to gain more control over how Reset works in in QuickStreams, Quick has gained two x one extension recently, which is ResetStreamAt. ResetStreamAt lets the sender define reliable boundaries, meaning, we set deliver everything up to that reliable boundary, everything up to that is protected, and after that, just throw the bytes away. And we are using that to avoid this partial work in two with two strategies. One strategy is we set this reliable boundary directly after the keyframe so that we can assure that at least the keyframe is being delivered to the to the viewer. And the other strategy we are proposing is we are protecting everything up to the last fully written object in our media stream. However, this breaks kind of breaks the promise that we already have been given with timeouts, because we are sending out stale data essentially with good intentions because we want to avoid partial work, but it's still stale data. So the question that we wanna ask, where should this reliable boundary be set? Where should it go? And we're trying to answer this question with three sub questions, what we which we are gonna look at next. One is where do when do get keyframes get lost? What is the actual latency cost of protecting the keyframes? And what does it actually mean for the viewer's playback? To test that, we have a it's a testbed with the publisher, a relay, and a subscriber. The publisher sends out Big Buck Bunny at 1.1 megabit per second and roughly 32 GOPs of roughly one second. And we have a relay in in the middle who wants who forwards the media and then applies the reset strategy. And behind a NetEm link, we have a subscriber which renders the incoming media, so we have a playback statistics. On the NetEm link, we apply certain conditions, A little bit of loss in the lossy condition, the 2% congested. We are limiting everything of the the connection to one megabit per second, so we are structurally congested. And the severe condition combines the 2% loss and the one megabits the one megabit constraint. So here is the ratio of keyframe number of keyframes actually delivered to to the viewer. What you see here is that the ResetStream loses keyframes, and the shorter the timeout is, the more keyframe it loses. And this is not even really tied to resources because the lossy condition has no restrictions on bandwidth whatsoever, so there is enough resources. However, the whole loss detection and retransmission time takes too long, and the keyframe is eventually lost. Where why? Whereas our result add strategies both are protecting every keyframe, and we can say, well, we have lost no keyframe due to due to a reset. So looking at latency. So there are three things that I want to point out here. One, our two reset add strategies are working almost, in this case, exactly the same. And this is the case for a lot of the cases that you're going to see or going to read in the paper. So the green no. The the blue and the yellow line are overlapping completely. Second, we have run a no timeout scenario, so where we are just plainly delivering fully reliable the key frames or the all the frames to the other viewer, and you can see here that it's accumulating latency over time. So in the end, the viewer is roughly ten milliseconds be up ten seconds behind the live edge. Whereas all three reset strategies are able to keep the delay to to manage the delay. What you can see between the reset strategies is that, judging from the graph, that the ResetStream is a little bit faster than the ResetStreamAt strategies. However, this is more or less a survivorship bias because ResetStream loses keyframes, and in general, it loses the keyframes which are harder to protect or harder to harder to transmit. Hence, if you are correcting for that on matching the keyframes, so keyframe matching the latency for keyframes delivered in the ResetStream and keyframes delivered in reset ResetStream at. The difference between the two melts down to zero to fifteen milliseconds, and it's equivalent statistically equivalent within fifty milliseconds in most of the conditions. So we can protect the keyframes at minimal keyframe at minimal latency cost. So when we look at playback starts, so we are monitoring what the viewer is actually rendering and everything well, whenever the player doesn't render anything for longer than a hundred milliseconds, we count it as a a as a stall, and we we cut the time. So you can see here that in our with our scheme, with our reset add scheme, we can limit the amount of freezing, essentially, the buffering, to one second. It's more or less one GOP size. This is because, well, we are delivering every keyframe, which means at least every second, the viewer has something new to render. However, and we are not fully we are not decreasing the overall stall time. Indeed, we are kind of increasing it. If you look here, the number of stalls and the total duration of the stalls, we are slightly above the the for the duration of the stalls, we are slightly above the ResetStream. We are highly highly above the ResetStream with the number of stalls, 3.2 versus 14 for our reset add strategies. However, well, as I said, we are able to limit the maximum amount. We are we are bounding the worst case in our scenario. So what you can see in the end is if if you protect the keyframes, you keep every every group of picture decodable. Protecting them has an as it is in our scenario, no significant latency cost. However, we are not producing fewer stalls, but we bound the worst case, meaning that you have the choice if you have one long stall or many shorter stalls. So and if you wanna if you keep the keyframe, you can protect you can bind the stall. Everything of my work is publicly on GitHub, so you can you can check out the the relay implementation. You can you can check out the test bed. You can check out the dataset. And for future work, we are planning something on relay scheduling and prioritization among testing different multimedia sources. And with that, I'm happy to answer any questions that you have.
[00:58:30] Gorry Fairhurst: Yeah, we can take some questions while people are coming. Did you consider different types of video, where the keyframe to predicted frames might be different in volume?
[00:58:46] Sebastian: Not on the paper. I did some private testing. It changes, but you have then to change the amount of bandwidth that you are using in your test batch because otherwise, it doesn't make sense. You cannot try to transmit, like, four k video over a one megabyte link or 56 k modem connection and expect anything but crap.
[00:59:13] Gorry Fairhurst: Yeah. I was just thinking Bigs Bucks Bunny probably compresses quite nicely compared to some moving scenes, etcetera.
[00:59:20] Sebastian: Yeah. Yes. That mostly that doesn't so much impact on the on the key frame size, but it impacts a lot of the p frame sizes. So that is the main difference here. So key p key frame is more or less constant because it's
[00:59:35] Dirk Kutscher: the whole picture,
[00:59:36] Sebastian: but keyframe depends heavily heavily on the scenery. Yeah.
[00:59:40] Sebastian: But yeah.
[00:59:41] Gorry Fairhurst: Can you
[00:59:42] Lucas Pardue: Lucas, quickly, the ResetStream draft progressed to IETF last call, like today, maybe. Thanks, Corey. So it's great to see research like this that's looking into the practical kind of outcomes of using it and the trade offs that people could make and that kind of thing. Just on the previous slide, you mentioned well, it's gone now. Don't worry. But something about prioritization scheduling. One of kind of the open areas in in a quick implementation is how to schedule retransmits of lost data.
[01:00:14] Sebastian: Funny you're saying that. During the hackathon, I I I implemented something where you could where I would prioritize fresh data over even retransmissions, as ex exactly for this case, didn't perform as bad as good as dropping the data altogether.
[01:00:36] Sebastian: Yep. I
[01:00:37] Lucas Pardue: I But but it's still very fresh. Yeah. I've got no more insight other than that's that's a cool area to keep looking into because implementers just don't know
[01:00:45] Sebastian: the answer. That's very much very high on my agenda.
[01:00:47] Francois Michel: Cool. Thank you very much.
[01:00:49] Sebastian: Thank you. Thank you.
[01:01:06] Suresh Krishnan: Should have control, Ali.
[01:01:07] Ali Begen: Sure. Actually, us. Okay.
[01:01:11] Gorry Fairhurst: Hold it if you like. Yeah.
[01:01:12] Ali Begen: Alright. Good afternoon. My name is Ali. Ali Begen from University. This is a joint work with several students of mine. I've been working in the streaming video space for almost two decades now. First on the HLS type of stuff, now more recently on Media over QUIC. And we have our own implementation called Mocktail, and this is one of the projects that we developed on top of that open source project. So today, I'm gonna talk about the watch party use case using Mock Transport. Watch party is something the young people, not necessarily the people in this room, are actually, you know, use it. It's a it's a it's a synchronized viewing of live video or on demand video. And at the same time, you have this good communication between the between the participants. So there's a video conferencing part, and then there is a streaming part of that content. And we are gonna use MoQ. And here, we are talking about a dual mode dual type transport approach. And here, we are referring to reliable streams as well as unreliable datagrams in Media over QUIC. And the source project and everything related to our open source is available at MoQtail.dev. So quick probably, don't need to introduce it too much, but a couple of important aspects that we use in this project. Obviously, there is prioritization between different quick streams. There is a user space implementation of congestion, content, rate control. It gives you different prioritization capabilities. You can discard stuff. You can cancel stuff without really harming the connection. You know, all those things are heavily used in Media over QUIC transport. And everything is encrypted. So if you are gonna use some network level quality of service, you need to use multiple quick connections, but that's not what what we are doing here. The previous presentation, thanks to Sebastian. He gave a quick introduction to MoQ transport. I have also one slide on this. So we started working on MoQ transport almost four years ago. Right? So it is one of the I mean, the one of the most important features of Mock Transport is that it is latency tunable. It's not something we had with WebRTC. Yes. We have it for Dash, but then that's, you know, multiple seconds of latency we are talking about. Here, we are talking about any latency between the near real time as well as tens of seconds if you want to. It is cache friendly, pops up distribution protocol, as well as it can be used for ingest. So we have publisher subscribers, and then we have relays in between that will provide the fin out trans fin out for for our content. Now the MoQtail implementation supports both transport to over web transport as well as over real quick. And we have all these media tracks and then the control streams back and forth between publishers and the subscribers. This is a chart from one of our recent survey papers. We can customize this, you know, classification as much as we would like to, but here is one. So we have high latency applications and a typical latency, low latency, and then ultra low latency and near real time latency applications. So what's really most important here is this low latency and ultra low latency part. So for many years, we have DashHLS for almost two decades where we can achieve tens of seconds of latency very easily. It is highly scalable. And then we developed low latency dash extensions. And then following that, Apple developed low latency extensions for HLS that will give you about you know, latency is about, you know, five seconds, six seconds, ten, seven seconds, and so on. And then we obviously have WebRTC for near real time or real time communications, but then there are a lot of hurdles. Right? I mean, most of the professional video entertainment video requires, you know, stuff like PRM and the and so on, and WebRTC doesn't allow you to do any of that. So, initially, we started with that mindset, and Media over QUIC transport would address this nice little gap, like a few seconds of latency, and that's it. But then we realized that, actually, once you have such a powerful protocol, you can implement the entire latency spectrum with Media over QUIC transport. So we don't need to do dash HLS for on demand applications and then WebRTC or something else for real time applications. You can do a single transport and get done with it. Now the use case we are looking in the in this paper is social core viewing, also watch party. I know there are other names like share play and so on. So, you know, it is really like a digital living room. Multiple people get together, and they are trying to share the emotional contact context. Right? And then this proves to actually increase improve the retention and engagement, so many companies are building services around this. Now, what's really important here is that whether we are seeing a live content like a football game from last night or or like you are watching a movie with your friends, right, You know, what's really important is that everybody sees pretty much the same frame at the same time. I mean, if everybody's at latency of ten seconds, that's still fine because you are still synchronized with your friends. Right? There are studies we have shown that anything less than the synchronization difference, anything less than two hundred milliseconds is really socially, you know, transparent. So it's like I mean, you feel like you are fully synchronized with your friends, right, coworkers. Anything less than half a second, it is okay. It's acceptable. But then beyond five hundred milliseconds, it becomes a spoiler. Like, you know, one of your friends is gonna see the goal, and then he's gonna cheer in the conferencing session. And then you haven't seen it, but you heard about the goal, and it's gonna you are gonna be quite miserable in that at that moment. Right? So we are trying to achieve this strict synchronization between the players, between the viewers actually, and trying to achieve also scalability as much as fidelity. So the quality of the stream is important as well. So we have done some tests, like our relay is in Frankfurt, Germany, and then we have we have players in Istanbul and then all the way on the West Coast, Vancouver. So even the round trip time is about two hundred to three hundred milliseconds, and we try to keep a synchronized playback between those two players with this protocol. So so far, most of the watch party implementations out there have have been implemented on using two different stacks. The LL-HLS for the streaming part and WebRTC for the video conferencing part. That means you need to maintain two different stacks at the same time. And not only that, you need to share the bandwidth in such a way that WebRTC doesn't starve the streaming media, and streaming media doesn't starve the real time media. Right? So if we do this in a single stack and then if we share the congestion state, actually, things become a lot easier to do, and we can we can control, you know, which streams are gonna use how much bandwidth. Right? So we are using instead of WebRTC, there's HLS Frankenstein. We are gonna use this, you know, MoQ transport, so we will have a context pipe where we are streaming the the the shared media. And now we have the excitement pipe, is gonna replace the WebRTCic component. So we have a single relay. I mean, obviously, we can have multiple relays that people are connected to, but then the single relay is gonna carry all the traffic for us. Right? So there's a it's consolidated infrastructure, first of all, and most importantly, it's shared congestion state through the quick section. And then it is payload agnostic. Mock transport doesn't care whether it's media or notification data, and everything is gonna be prioritized based on the publisher or the subscriber and so on. Oops. Yep. Here is here is the architecture. So we have the main content. Maybe it's Big Buck Bunny or the Spain versus Argentina game. Right? It goes through the live encoder or on demand encoder. It gets to the publisher, which is our also our our own publisher. And then we have a relay. Again, we can have multiple MoQ transport relays. This is a single peak connection that is transmitting multiple tracks, like audio track, video track. And then if you have rate adaptive streams, there could be multiple video tracks. And then we have timeline track. We have catalog track and then all the other stuff that you need for streaming media. And then between each part here, it could be, like, two guys or 20 guys. Doesn't matter. They have open single connections, quick connections to to the relay, and then they will be sending their audio video and the text messages, and then we'll be receiving the audio video from other users in the in the party. So the dual transport dual pipe transport approach is as follows. We have the context pipe, and then we have the excitement pipe. And then there's an excitement pipe for every other user. So if there are 10 users in the system, then then there will be nine such pipes. Right? So things are going to be prioritized based on their importance. First of all, we wanna make sure that, you know, instead of really fully making sure that your friend's video is gonna come through to your computer or mobile device, we wanna make sure that the content you would like to watch is actually flawless. So the context pipe is always higher priority than the excitement pipe. I'm gonna show an example here, hopefully, if the slide switches. Yes. So the main tone content comes in. Let's say that's the football game. And from each other user, you are getting an excitement path pipe. Right? So the control stream is the highest priority. That's based on the MoQ transport. And then the next high highest priority is the main contact, the football game, audio video tracks. And then from each part here, we have the party audio as well as we have party video. Now the party audio is medium priority. I would like to hear from my friends, right, before I actually see them. That's really I mean, more important than the video. This is latency critical, and but it is unreliable. I can transmit them over quick datagrams. When I look at the party video, it is lowest priority. So if I have still have bandwidth, I would like to use that. And it is latency sensitive. It's cancellable. Meaning that, you know, if it's gonna be late, we cancel that at the publisher or at the relay. Just like the previous presentation showed, we call it early discard. So if you have less bandwidth, the video is gonna be dropped. If you have even less much less bandwidth, then the audio is gonna be dropped as well, but making sure that the football game is still streamed successfully.
[01:12:45] Gorry Fairhurst: So
[01:12:47] Ali Begen: if you wanna test this out, it is a watchpartywatchparty.MoQtail.dev. We code and then everything, the papers representations are already on our website. I I encourage you to look at other demos as well. We have different demos implemented at different versions. Just, hopefully, in a couple of days, we will be version 18 compatible, and we will be, you know, participate in these hackathons. So any questions for now? Thank you.
[01:13:27] Matthijs Engelbart: Questions?
[01:13:29] Ali Begen: Okay.
[01:13:32] Suresh Krishnan: If nobody has a question, I have one. So you talked about the Frankenstein solution with the WebRTC dash. Do you have any numerical I understand the subjective thing why the complexity hurts. Right? But do you have any results showing that, like, numerically, this is superior?
[01:13:46] Ali Begen: Do I have any what?
[01:13:47] Dirk Kutscher: Results. Results. Numerical results. Yeah. I
[01:13:51] Ali Begen: we have results in the paper, and, you know, I wanted to cover them as well. But then considering the twelve minute limitation, I didn't want to bother you with those results. Please look at the results in the paper. And we are I mean, those are experimental results, not just the simulation results. So I hopefully, they will give you a good understanding of what's happening.
[01:14:12] Suresh Krishnan: Thank you. Thanks.
[01:14:13] Ali Begen: Thank you.
[01:14:48] Presenter: Alright, awesome. Hey everyone. I'm very happy to talk about my paper about Multipath QUIC scheduling. We are on the transport layer here, so Multipath QUIC, MP-QUIC, that kind of stuff. First of all, I would like to go into the problems. Why do we need, or in our opinion at least, the new Multipath QUIC scheduling model? So we found that with MPTCP, for example think about BLAST, ECF, schedulers are defined and implemented as per packet decision functions. That means they get the pre assembled packet and ask themselves the question which path should this packet take and make a decision. That means that the scheduler can only decide where this packet is supposed to go, but not what should be inside that packet. If the scheduler wants to do additional logic such as to duplicate data across multiple path or reinject in flight data, it needs to hook its logic into different parts of the protocol implementation. So problem one is a very tight coupling between the scheduler and the protocol implementation. Now coming to Multipath Quick, we found that compared to MPTCP, it has a lot of additional scheduling facets. The first one of them is in addition to reliable data, we can now also send data unreliable. And the scheduler will want to schedule datagrams different to stream frames. Second, we have data multiplexing. So we have multiple streams and those often also have different priorities which the scheduler wants to take into account. And finally, in QUIC control data such as acknowledgments and flow control updates is carried explicitly in QUIC frames and not compared to TCP implicitly in every TCP packet header. And the scheduler can make scheduling decisions based on the core for the control frames. So the per packet scheduling model is really limiting the possibilities of MP Quick. And the third problem, we find that for an MP-QUIC scheduler to be really effective, it should be really tightly coupled with the application. So the application should have a lot of control over how the scheduler works. With MPTCP, this is difficult because it's implemented in the kernel and you need to or you need to work with interface the kernel gives you. MP-QUIC, however, or QUIC in general, has a lot of potential for type coupling between application and and to end transport, but we find that currently this isn't realized. That's why we set out to design a better interface for MP Quick Scheduling according to the following or with the following goals. First, we want to have a uniform inter interface that supports all these facets that I've just mentioned. Second, we wanna encapsulate the the scheduler to prevent this tight coupling between scheduler and library. And third, we want to enable the application driven scheduling. Towards this goal, we make the following three design choices. First, in our model, the scheduler runs upfront. That means the scheduler runs once before packets are sent. And in the single core, it makes all the scheduling decisions. And it's inspectable. What we mean by that is that the scheduling decisions are actually expressed as a persistent data structure and not as the result of repeated function calls. Second, while the scheduler has a lot of power, we want to make sure that the protocol still behaves according to its specifications, so it hears flow control limits and congestion control limits, for example. And third, to enable the application scheduler coupling, we make the scheduler application owned. That way, the application has the most direct access to the scheduler possible. Next, I want to explain our model. We introduced five new components for it. First one is the const state object, short for connection state. It's a snapshot of the quick connection state, so it includes information about the open streams, how much data is queued, but also about the path, the congestion window, round trip time, and it informs the scheduling decision. Additionally, it has a list of the pending control frames. The pending control frames are the frames that need to be sent right now. The assembly of this object is triggered by a new egg arriving from the peer, by the application writing new data or by a timeout such as a PTO file. ConState is the input to the scheduler. The scheduler makes the scheduling decision and persists it in per path queues, of which there are two. First, there on each path, there's a control frame queue. The scheduler looks at the pending control frames and assigns them to the path on which those shall be sent. The second type of queue is the data token queue. Here, we store the instructions for which data should be generated for this path. It is subdivided into different levels, which determine the order in the tokens are pro processed. There are two types of tokens: stream tokens, which correspond to stream frames, and datagram tokens, which correspond to datagrams. And they have a certain number of parameters that I will go into next. After the scheduler has completed, the packet assembly loop runs. This drains tokens from the per path queues as long as packets can be sent, which most of the time means as long as congestion window is available on any path. I think this model can be best understood by going through an example. And I will construct stream-aware minRTT scheduler next. We have two high priority streams and a lower priority stream. And the full scheduler can be shown by looking at the multilevel token queues. The stream-aware minRTT scheduler is all about processing order. On one level, minRTT scheduling. That means in which order should I generate data for the path? I want to generate data for the low latency path first before others. And on the other level, stream-aware scheduling. That means I want to send high priority streams before lower priority streams. The levels, we have the central rule that lower levels are processed first. And we can express both of these ordering constraints through levels. Each level contains a number of tokens. In our case, we only send stream tokens, and they have three different fields. The first one is the stream ID. Token with a certain stream ID generates a stream frame for that stream. The second field is the offset. There are different parameters. Here we choose the unspecified offset. What that means is when that token is processed, the library looks at the stream sent queue. If there are retransmissions queued, first the retransmissions will be sent. Afterwards, the next on sent offset will be used. And finally, the length field. This one determines if the token is requeued at its level. And for the minRTT scheduler, we use the per system length. This means that the token will be requeued as long as the stream lives. And finally, want to stress this once more, the queues are persistent. That means the scheduler installs this configuration once when it is invoked first, and it will only change it if new streams are opened or if order, if the RTT order of the path changes. Now let's jump into this example. I have here some example numbers. These are the number of bytes queued for the streams for sending and how much congestion window there's available on the two path. The central rule is the front token from the lowest available level is processed first. That means, in the first step, we process the stream 10 token and send a stream 10 frame on path zero. This stream 10 frame will fill all space that is available in the packet after the control frames have been added to the packet. Next, it will be re queued at the end of its level because of the persistent length parameter. And we will produce a stream two frame. Now we don't have any data left queued for stream two, so the token will be disabled. The next step, again, we will create a stream 10 frame, and now we don't have any congestion window available on path zero anymore, so we will disable the entire path. And due to this pro processing rule, proceed with level three on path one. There we produce a stream 10 frame on path one. And now since also path one has no congestion window available anymore, both paths are disabled. At this point, the pack generation stops, since we can't generate any packets anymore and quick sleeps until the next trigger, like for example an acknowledgment arrives and the constable will be assembled new and the procedure starts again. So at this point, you might be asking yourself this. So Minot, here's the simplest multipath scheduler. You can imagine that this isn't simple to understand. It's much more complex. What's the point? And I would argue that the point is that with this token syntax, we can make very small changes, but very powerful changes at the same time. I have two examples. So we can expand the MinRTT scheduler with redundant retransmissions. So think a length from o to o plus n is lost. The scheduler can additionally queue tokens on all available path with the specific offset o and the length n. Because we recruit on level zero, those will be processed first. And because n is a concrete number of bytes, the token would be dropped when those bytes have been sent. The second example is we can actually turn the full minRTT schedule into a scheduler by changing the offset to a specific number. In this case, first zero. What this does is each time the token is processed, n bytes are emitted, n is how many bytes fit into the packet, and the token is re queued with offset increased by those number of bytes. That means the token will go through the full stream from zero to its end. However, ACK bytes are skipped so that no obviously wasteful data is sent and because most implementations don't queue that data anymore at that point. And the final point I want to make, we can also schedule control frames, which is something that I've not really seen with any other multiple scheduling approach. We can, for example, schedule frames on the faster, faster path to make x arrive soon. I will go quickly over the evaluation, what I want to tell you is that we actually implemented that and that the implementation is publicly available, and that we conducted an experiment where we used x scheduling policy to improve the stream completion time. A quick conclusion. In my opinion, the big strengths are the clean encapsulation of the sketching logic, that we have the potential to actually move the same schedule definition between different QUIC implementations, and that the data structure allows to inspect the scheduling policy, which eases development and testing. The main limitations are that we have no control over packetization, because we can only control which frames go where, and that we can't choose when the logic runs. Thanks.
[01:28:24] Suresh Krishnan: Thank you. Any questions?
[01:28:38] Gorry Fairhurst: I'd just like to say thank you for doing this because we've talked so much about how to do scheduling and this is kind of a really interesting approach to it. So thanks ever so much for bringing it here.
[01:29:14] Sebastian: K. Can you hear me?
[01:29:15] Suresh Krishnan: Yeah. I can. Sebastian, you have slide control. Check if you can move the slides. Use your right arrow, left arrow. Okay. You're good?
[01:29:24] Sebastian: I can. Yep. Thank you. Cool. Thank you. Hi, everyone. I'm Sebastian from Ericsson Research. I'm presenting our paper quantifying QoE-Aware Resource sharing potential under time variant spatial complexity. This is a measurement and analysis study. The one sentence to remember is that for real time interactive video, equal bitrate is a poor proxy for equal quality. And we quantify how much dynamic quality of allocation could gain. So, yep, first, the setting, we care about real time interactive video. For example, cloud gaming, remote rendered XR, remote control of drones or machinery. The common thread there is that there is a human or AI in the loop with a finger photon latency budget of roughly thirty to one hundred milliseconds. And there are two consequences here. The video must be encoded in real time, and there is no room for the long buffering that on demand streaming relies on so that usual trash tricks do not apply here. Quality Yep. Of experience for this traffic has three parts spatial, how good the picture looks, the temporal quality, the high end consistent frame rate, and input lag quality, which is the finger to photon delay. This talk focuses deliberately on spatial quality measured with VMAF. That is a scope choice, not a claim that the others do not matter. So our allocation works at one second granularity on aggregate rates. So it doesn't, that frame level delay and the temporal and input like components, anyway, need packet level evaluation, which is our future work. So when I say quality in the rest of the talk, I mean spatial picture quality. So some framing for this audience. Today, the network typically shares capacity by flow rate fairness, roughly equal bit rate. And this call argued back in 2007 that equal rates are a really poor proxy for equal benefit. So we make that concrete for real time interactive video. Given two sessions of the same bit rate, you get very different picture quality because the spatial complexity differs. So the question of the talk is, if we allocate for quality instead of equal rate, how much can we gain? So this is, again, a measurement and analysis study, and we quantify the upper bound of the gain. We don't claim that this gain is achievable, and we don't have an algorithm which can actually achieve that gain. So this is the object we work with. This is a spatial complexity curve, or SCC. It maps the encoding bit rate on the x axis to the picture quality VMAF on the y axis. And we introduced it actually in our ANRW twenty twenty four paper two years ago. There are three anchors there. VMAF 50 is the minimally acceptable quality. 70 is good, and 90 is excellent. And as you can see on the picture, different content needs very different rates for the same quality. Look at, for example, the 70 good QE. The spatially simple card game, the yellow curves, which is good quality below one megabit. The spatially complex third person game needs several megabits for the same quality. And there are two kinds of the curves. Look at them. The solid lines are the average complexity of a whole clip, which is about one hour of gameplay. And the dotted lines are the average complexity of individual two twenty seconds cut from the same clip. And for example, the blue arrow shows the key point. Even sessions from the same one clip spread widely. So per clip average already hides a lot. And as we see on the next slide, per session average is also not enough. That's, again, the dotted lines. So first, the bad news, and this is one of our central measurements. One session, one quality targeting encoder, four VMAF target levels. Remember the per session average from the previous slide? So now we are looking inside one of those sessions. The rate needed to hold VMUF 70, for example, swings widely within a two twenty second session. In the simple early phase, it sits well under one megabit when the scene turns complex, around 100 in. The same targets need several megabits. So there is an order of magnitude difference within one session. So a single average per session rate cannot describe it. And this within session time variance is our first contribution. And it is basically as large as the variation across different content types. The consequence is simple. No fixed rate can track track this. So if you size it for the average, the complex seconds will fall with acceptable. If you size it for the peak, you base capacity most of the time. Actually, that's what we basically have to do most of the time if you are doing rate based traffic management here. We have to size for the peak. Yeah. But but now the good news, this is the aggregate demand of 30 quality target sessions and one line per target VMAF. So compare the small relative variation here to the wide single session variation on the previous slide. So you can see that the sum is much smoother. For example, the VMS 70 aggregate, you can see that there is much, much smaller variation for aggregate. And the reason is simple statistical multiplexing. There is no reason for the complex scenes of independent sessions to line up in time so their peaks are uncorrelated to that bridge out. And two things follow from that. The aggregate is much more predictable, which helps provisioning the link. And at any instance, some sessions are simple, and some are complex. Simple have some spare capacity that could cover the complex ones. And the second point is exactly why a dynamic second by second reallocation can exploit this and what the static per session rate cannot work. And here, we model dynamic session second by second reallocation. There can be other solutions to approach this game. So how should we allocate? We make the policy explicit. It is a utility function over quality. Every VMAF value is versus certain utility, and we allocate to maximize the total utility achieved. There are three end goals here. Again, VMAF 50 is acceptable. It is versus utility 100. VMAF 70 is good versus utility 120. And VMAF 90 is excellent versus 130. And below 50, they go zero. The shape encodes two design choices. The first is a big jump at 50. That is kind of admission control. The session below acceptable quality is worth zero utility. In practice, we drop that session rather than let everyone let it drag everyone down. So the big step stays, get every session to acceptable first. And the second choice is that the slope decreases. So each step up is worse less than one than the one before. So concretely admitting one new session versus as much as upgrading five sessions from VMA 50 to 70. We call it breadth over depth, keep many sessions acceptable, rather than you could make a few excellent. So why is this the right policy there? Because any session can become complex at any second. So suppose we spend capacity making the simple sessions excellent, then there wouldn't be capacity left for the spiking flows. And those would fall below acceptable. So breadth first avoids that. It always gives the next bit to the session nearest to the threshold so nobody falls out. That is exactly what today's rate fairness cannot do. This one also keeps the e sessions quality consistent over time, which is what long sessions like you ultimately need and long session QE ultimately needs. One more point about the curve, it's fixed in VMAF, but the rate needed to reach a given VMAF changes every second, as we have seen. So the rate to utility mapping is time variant. And that is why the allocation rebalances on its own as the spatial complexity moves. So finally, one cover to keep in mind for the rest of the talk, everything from here is Oracle. Perfect per second complexity for every session. Already seen for the future, and this is deliberate. We are not proposing a scheduler. We are not proposing a solution. We are measuring the ceiling of what is possible. Closing the gap to and having practical optimization is follow-up work. So read every result after this as an upper bound. So a simple example of a single second allocation, three megabit capacity, five seconds. Just count how many bars clear the acceptable line at VMA 50. The allocated rate is written inside each bar. So for equal VMA from the left, it equalizes quality. Every session sits about VMA 46, just below acceptable. Zero acceptable, zero utility. Equality, but everyone loses. Rate fare in the middle speeds capacity evenly. Every session gets 0.6 megabits. So only the simple card game clears the bar. One out of five, again, rate, but VMAF ranges from 36 to 79, purely for the spatial complexity variation. And max utility, according to the rules previously, and utilization maximization on the right drops the one hopeless session, the red bar, to zero, and uses the capacity to lift the other four above 50. So four out of five acceptable. So here's the punch line. Maximum utility at equal-VMAF has about the same average VMAF, roughly 47, yet their average utility is 86 versus zero. So about same average quality, completely different average utility. And average VMF hides most of the things. What matters is how many sessions clear the acceptable line according to our definition, of course. But in general, keeping consistent quality is important. Below that line, the picture is simply not usable or very bad. So the same average VMAF can mean either no session is above the line or most of them are. Okay. So that was a very, very simple example. Now the behavior over time at 50 megabit per second, but, again, shared by thirty seconds, two figures, one trade off. On the left, you can see the per second minimum VMAF, the first session at each second. And you can see that both maximum utility and equal-VMAF holds that minimum above the floor of 50. Rate fare, which is today's practice, repeatedly collapses there. And on the right, can see the average VMAF. And here's the catch. Equal protected the minimum best, but it has the lowest average. It drags everyone down to pull the worst session up the floor, while max utility have the highest average and still holds the floor. So So the tradeoff is equally must hold the worst session up by pulling the good ones down, while max utility protects the worst but keeps the typical quality high. And in most cases, this makes more sense. So this is the one of the main results of the talk, how much bottleneck capacity you need to keep everyone acceptable, dynamic versus static. On the y axis log scale, the fraction of session seconds below acceptable quality. Along the x axis, we increase the bottleneck capacity, C. And you can see that the two dynamic method, max utility and equal-VMAF, both using a one second oracle drive that fraction to zero and thirty five megabits. Every session, every second acceptable. And the static methods never get there, not even perfect, at 100 megabits per second. And that is with a perfect oracle for each session's average complexity. So it's a different oracle. The rest would be called maximum utility per click clip. So the gap is roughly a factor of three, thirty five megabit per second dynamic versus more than 100 for static, more than 100 for allocating for the peak for something like rate fairness. And the reason is fundamental. Fixed rate for average under serves the complex session. And if we size the rate for the peak, it will base capacity. And if we can apply second by second reallocation, it works because at any instance, some sessions are simple and their spare capacity can serve the complex ones. And that is why static waste is too much. Today's whole pipeline, encoder, congestion control, and rate for sharing is built around the peak of each session, as if every session were complex all the time. And closing that gap is a networking problem as much as a codec one, and that is why it should matter for this audience. So to wrap up, three points. First, the spatial complexity is time variant within a single session as much as it varies across content types. Second, the utility functions makes for the allocation policy explicit and tunable, and it beats the rigid allocation. And third, the potential is large. Dynamic methods keeps all 30 secondtions acceptable at 35 megabits, while no static method does that even at 100. And one method, the Oracle is deliberately an upper bound. Closing the gap to a practical estimation is an open question, and it connects to our QoE-Aware Resource Sharing work on the in session control loops. Yeah, the next slide has some pointers about further reading, and I'm happy to discuss any of these and open for questions. Thank you.
[01:43:37] Gorry Fairhurst: Are there any questions from the room? Okay. In that case, thank you for an interesting talk and for, you know, bringing lots of data here. It's always nice to see lots of data. So thank you.
[01:43:57] Sebastian: Thank you.
[01:44:16] Ingrid Tovin: Alright. Hi, everyone. My name is Ingrid Tovin. I'm here to present cTAPS, which is an implementation of IETF transport services over QUIC TCP and UDP and has been my master thesis at the University of Oslo for the last year. So first of all, what is transport services? Well, transport services is a new way of thinking about the transport layer. Instead of binding tightly to a specific transport layer protocol, you instead have this generic API which then handles both protocol selection, but also lower level network selections for you. For example, setting the DSCP value in the IP header. cTAPS itself is naturally, it's a C implementation of transport services and it's open source and has QUIC support. And it's the first open source implementation of transport services which supports QUIC. So since QUIC was a large motivation behind developing cTAPS, we also performed a survey of RC9000 and found that 24 out of the 28 surveyed features I found could be either fully or partially supported through the TAPS spec. However, not all of these are supported in cTAPS and that is primarily just a limitation of the implementation itself. So here is a example of a important feature of QUIC that is, available through this, generic API in cTAPS. So here you have a quick normal quick client through picoquic, and you have a cTAPS client who and they are both able to recover after a local path event where you can see that the equivalent TCP connection just fails. So the strengths of a generic API like this is that you get a lot of features such as security, protocol racing, DNS, etcetera. And you get them all through this single unified interface. And this also enables wider protocol adoption in that you don't have to redesign your entire application in order to run over QUIC if it's already compatible with TCP. It also makes use of easier use of network level choices choices such as DSCP by having generic configurations. You can set it you can set the configuration to be that you want the low latency con connection, and then the API itself will then make whatever choices it can in order to achieve that, which naturally would include setting the DSCP bit. It is also callback based by default, so it reflects how people actually use the network and it's not blocking like a socket would be by default. And for the limitations itself, it's cTAPS and especially a fully feature complete transport services library would be a fairly complicated piece of software and it's also fairly opinionated about exactly how you would design your application. Then minimum complexity is slightly higher than the simplest possible socket clients, but when you get to more complicated applications, the total complexity is slightly much lower. Transfer services does also have some gaps in its quick support. For example, it's currently not possible to open a unidirectional stream if the overall connection is bidirectional. But I will discuss this further in a talk with the quick working group on Wednesday. If you're especially interested, you can find the source code and my master thesis on GitHub. Thank you.
[01:48:34] Suresh Krishnan: Thank you very much. Questions?
[01:48:37] Sebastian: Go ahead.
[01:48:38] Speaker 16: Sorry. I didn't raise my hand. Hello. Michaud from TU Eindhoven. So I have a question regarding, like, implementation and the API itself, maybe more on the details. We all probably know the meme. There are 11 standards. Let's create a unifying one. There are 12 standards. So so how does this bridge this gap? Like, what is the way that it is in kernel? Is it in user space? How do you use it? Why? You see my question. Right?
[01:49:10] Ingrid Tovin: Yeah. I think, at least from my point of view, the strength of this is that it is in user space, and it's it only needs to be used by the wire itself. What's transferred over the wire looks exactly the same as if you're doing normal socket programming. So this is not something that needs to get adopted by both servers and clients. This makes sim yeah. A single application's life easier that way. So I think that's the biggest strength that it has there's no cooperation needed for this to be beneficial that way.
[01:49:45] Speaker 16: So it's only on the server side?
[01:49:48] Ingrid Tovin: You could do you can do both servers and clients, but this is more of a it's slight like abstraction library above sockets, so you don't have to bind tightly to the sockets themselves. But underneath, at least in cTAPS, I'm using sockets under the hood. Yeah.
[01:50:06] Speaker 16: Oh, okay. I see. Yeah. Okay. I that's it. Thank you.
[01:50:11] Suresh Krishnan: Thank you. Thomas?
[01:50:12] Thomas Witt: Thomas Witt. High level user APIs in networking and also in frameworks are a way to undermine standards by vendors and by application providers, operating system providers. So how would you position your work as a standard I mean, implementation of a standardized approach to to counter these trends over the last ten, fifteen years?
[01:50:42] Ingrid Tovin: Yeah. I guess at least from the very beginning, it is set of RFCs that this has been so cTAPS is the implementation of the transport services RFCs, and that is at least an attempt to open source and democratize this away from maybe a closed source software library that would be developed purely by Microsoft in comparison, for example.
[01:51:08] Suresh Krishnan: Thank you. Thank you very much for your presentation, and thanks for the questions.
[01:51:15] Gorry Fairhurst: Yep. Thank you.
[01:51:21] Suresh Krishnan: And with that, we come to the end of the program. So thank you very much Gorry and all the session chairs before Matthijs. Also, I'd like to just like thank you very much for doing it. And thanks to our reviewers. Like, you you are, like, amazing in everything. So we got all the reviews. We got, like, really, really good reviews and everything. Thanks for everybody for being an amazing audience. And, Thomas, if you come out Thomas is We
[01:51:48] Thomas Witt: also thank the speakers.
[01:51:50] Suresh Krishnan: Oh, yeah. Thank you very much to the speakers. Oh, they get to speak, Thomas. That's the thanks. But no. I'm just kidding. Yeah. Thank you. And everybody, like, really kept the time and the discussions are, like, amazing. Thanks for being a wonderful audience. And a few and thanks for Dirk for trusting Thomas and I to do this work and, like, giving us the chance to do this. Also, few things. The the program page on the site is, like, requires updating. So we have the slides on the data tracker, we'll add the links to it by end of day tomorrow. We also got the proceedings link from ACM today, like, during the meeting. So we'll update the links to the papers there as well and also the video link. So if you check out by the latest by the end of the week, you'll have all the things in the in the program page itself. And, Thomas, do you wanna add anything come come up? And yeah. Okay. Dirk?
[01:52:50] Dirk Kutscher: A Yeah. Quick comment. I just wanted to thank you, Suresh and Thomas, doing this. It was really another highlight of the ITF. Thanks very much.
[01:53:11] Suresh Krishnan: With that, you get six minutes back to go for early dinner and end the schnitzels here.
[01:53:22] Gorry Fairhurst: Very smoothie. Very smoothie. Yeah. Good talks. Good momentum. Yeah.
[01:53:29] Dirk Kutscher: That wasn't me. Okay. It was them.
[01:53:31] Suresh Krishnan: I I found a vegan Viener schnitzel. Alright. It does fantastic. Okay. If you're
[01:53:37] Ingrid Tovin: into Well, you're you're vegan. Right? Or or vegetarian?
[01:53:40] Gorry Fairhurst: I'm yeah. I only eat fish and sometimes chicken, but I predominantly eat that. You do?
[01:53:46] Suresh Krishnan: Yeah. So the yeah. It's like a place called Dalani. I can send you a link. Yes. It's like really, really good. Yeah. I'm gonna watch, and I can go there. I'll be good. Yeah. Yeah.
[01:53:57] Gorry Fairhurst: Have exceptions to the rule, which is confusing, but,
[01:54:00] Suresh Krishnan: generally, like, 95% of the time.
[01:54:03] Dirk Kutscher: Sounds good.