Markdown Version

Session Date/Time: 23 Jul 2026 12:00

[00:00:08] Dirk Kutscher: Okay. Marwan is here, then we can actually start. Welcome. This is IRTF Open, the open session of the Internet Research Task Force. I'm Dirk Kutscher. I'm the IRTF chair. It's great to see you. Before we start, so it's Thursday. You have seen this a couple of times, but just make sure that we follow IPR rules of the ITF here. And so if you contribute or see any IPR relevant material, you are expected to let us know in a short time frame. This meeting is audio and video recorded. Please be aware. Also, have a look at the node well and this privacy and code of conduct. So in addition to what the IITF has, has a dedicated code of conduct document. So that's RC9775, which has some additional advice for ethics and science and so on. If you are in the room, please sign up with MeetEco. Scan the QR code or use the on-site tool that is linked from the ITF agenda. And yeah, so I'm sure that there must be some newcomers here to this meeting. And just to inform everybody, here in the Internet Research Task Force, we're not doing standards. So we are only doing research. And that means sometimes people confuse these things because the groups have similar formats and so on. But nothing we do here is actually creating standards. Some groups may produce some documents, some ROCs, but these are only informational or experimental. And so we have quite a bit of freedom in the way we do the research, and so it's not really also constrained by ITF processes in general. Okay. So we have 16 research groups in the IRTF at the moment, and 14 of those are actually meeting this week. That's great. So it's great to see so many of you. We have this model when we have new groups that we kind of start on as so called proposed research groups. And then we kind of see, you know, how things develop, whether there's sufficient community involvement, and so on. So we did this for the sustain research group and also for the space research group, but so they are both really successful, and so I promoted them to regular groups this week. So all of these groups are regular research groups. A quick note. So our groups here are sometimes, you know, also engaging with other communities or they do some outreach. We don't have to hold our meetings only here co located to the IETF. So that's something that I generally encourage. You know, also think about other formats. And so here's one. So on the topic of networking for AI, we will have a second workshop at Conext in December this year. So this is actually from from last year's workshop. So we had it in in in in Hong Kong. And then this year, there will be another one at CO NEXT twenty six. We and the iTAF is also managing the diversity travel grants that brings new people from diverse backgrounds to these meetings. And this time, we have been lucky to have four newcomers coming to our meeting. And we have already made the call for the next ITF in San Francisco. And we will open the call for ITF one hundred and twenty eight on Kuala Lumpur in late November. So if you are interested in this program, please pay attention. We will announce it on IRTF announce and our social media channels. We also had the applied networking research workshop this year this time on Monday here in Vienna. So it was a very successful workshop, and lots of really cool papers were submitted, and a good program was selected. We can so the co chairs were Thomas Schmidt and Suraj Krishnan. You can find all the papers and presentations on on the IRTF website. And, yeah, we would of course, we do this every year. So you can expect the call for papers for the next one early next year. And, so I'm now coming to the highlight of the session. So we are running the Applied Networking Research Prize, ANRP, with support from ISOC, Comcast, and NBCUniversal. And so the idea is that we recognize the best work in applied networking every year. So there is a nomination process and then a selection process where we select the best papers that was published roughly in the previous year time frame. And so the idea is to bring this excellent work to here, give people a chance to to come, experience the IETF, and, yeah, maybe also find new collaborations and ways to contribute to the IITF in the future. And so this is always bit of work in, you know, reviewing and selecting these papers. So we had a really good selection committee. Thanks everybody who helped with that. Today we have two very exciting talks. I'm really happy here to have Rumaisa Habib to talk about her work on formalizing dependence on web infrastructure. And then I have Diwen Xue here, and he will talk about fingerprinting deep packet inspection devices. And just for orientation, so we have these two talks. There's enough time for discussion, so don't hold back your questions. And then afterward after that, I will give you a quick update of what just happened in a side meeting on networking for AI. And, Rumaisa, you can already come up, and I start introducing you. So, Rumaisa is a fourth year PhD candidate under the supervision of, Sakia Domenic in the, empirical security research group at Stanford University. And her primary research focus is on large scale measurement studies to understand the web beyond the western context and building systems to allow researchers conduct these studies more efficiently. And I have to stop sharing, and you have to just a moment. Got it. Okay. Welcome, Romesa.

[00:08:58] Rumaisa Habib: Alright. Thank you so much for the introduction. Hi, folks. I'm Rumaisa. I'm a fourth year PhD student in the Empirical Security Research Group at Stanford. This is a paper that was published last year at SIGCOM, and I had the pleasure of working on this with Kimberly Gautam and my adviser, Zakir. The focus of our work is web infrastructure centralization. I'm talking specifically about the providers that make websites reachable, reliable, and trusted, so hosting providers, DNS providers, and certificate authorities, more generally, the concentration of power on the web. And this is not a very new concern. It's been a topic of conversation at the IETF before. In particular, RFC nine five one eight talks about the growing concern of power imbalance on the web while also pointing to some natural positive reasons why this is occurring. We don't argue for either side in this work, but it's clear that this is something that the standards community cares about. In fact, there's a dedicated working group at the ITF, DINRG, which I think you are the director of, which looks at this subject. They focus on defining and characterizing centralization as well as understanding its ramifications. We'll come back to these objectives in a bit. In the academic world, several recent papers have done an excellent job showing that web infrastructure is becoming increasingly centralized. So given all this conversation about it in both the standards community and the academic community, it's important to nail down exactly what we mean by centralization. However, all of these prior works have a varying understanding of what it means for a region to be centralized. They all agree on this informal definition of centralization, which is the concentration of an Internet function on a small number of providers. While seemingly descriptive, where they differ is what they use to quantify it. For example, how much concentration on how few providers would entail a highly centralized region? But also providers don't exist in a vacuum. What is the regional context of these providers? And this is where our work comes in. Conversations about dependence have largely centered around the centralization of huge hyperscalers, but we argue for a more systematic view that not only includes concentration of technologies or companies, but the concentration of regional power. This shifts the conversation to look at real lived experiences rather than just the power that the cloud players and Amazons of the world have. While centralization is a lot about these large companies, it is also about the majorities that only appear when we look at the finer details. To aid in this analysis, we also provide statistical metrics to continue monitoring it longitudinally. This matters for tracking progress on centralization or understanding how realistic and fair a protocol deployment would be. Revisiting our problem start statement from earlier, our work makes two concrete contributions. We expand the conversation of centralization to include dependence as a whole and provide a suite of metrics to continue empirically monitoring it. Importantly, we focus on three of the objectives as they per se pertain to the web. We measure, characterize, and quantify not only centralization, but this broader notion of dependence on the web. A large part of the conversation on dependence is about centralization. Prior work understands it by looking at simple metrics such as the top end providers in a region and then the percentage share that they hold. While this does give a high level understanding that centralization is occurring, it paints an incomplete view when comparing between regions. So, for example, we found that the top five providers in Hong Kong and Azerbaijan both correspond to 59% of the websites they hosted, which may lead you to believe they're equally centralized. However, the distribution between and within the top providers is substantially different. So Hong Kong's top two providers cover 33 and then 12%, while Azerbaijan's covers 425% of market share, a much steeper drop off than Hong Kong. No single value of n would make this an even comparison. Like, you could try top five, top 10, but then juggling several of them makes this problem even more uninterpretable. Similarly, things get even more complicated when two regions have a different number of providers. With a distribution like this one, how does one even quantify centralization? Like, country b has more providers, but the top provider hosts more websites than countries a's top provider. So what gives? The key reason these top end metrics fundamentally fail is that they don't account for different distributions. They miss out on this nuance of varying numbers of providers and a long tail of globally unimportant but regionally important providers, which also leads to this other concept we introduced in this work, and is eek it is equally important in understanding dependence, regionalization, or the geopolitical dependencies between or within regions. Stepping back to the RFC mentioned at the start, even a seemingly decentralized ecosystem in terms of providers could be centralized on a region. So for example, a region could have could be depending on many providers, but if they're all located in a specific region, that region still holds power. Looking closely at regionalization lets us extract this broader context about dependence. These two concepts, centralization and regionalization, go hand in hand when describing dependence. In the following slides, I'll describe the statistical metrics we use to measure these concepts. Starting off with centralization. It was important for us to account for a varying number of providers as described earlier and not make any assumptions about its distribution. In in addition, we wanted it to be independent of specific providers and match a human level understanding of what centralization looks like. For our purposes, we borrow Earth's movers distance, EMD for short. Its definition is the minimum amount of work needed to transform one distribution into another. In our instantiation of this metric, we measure how far the distribution of providers in a region is compared to a fully uniform one, so one where every website is served by one provider. So remember those two sample countries from earlier. We can use EMD to reliably compare centralization between them even though they have a different number of providers without relying on a single snapshot of maybe the top two or top five. Importantly, though, we're not saying that this flat, fully uniform distribution is ideal or even possible in today's world. This is just something to give us a baseline by which we can compare the skew of distributions. So it's meant more as, like, a comparison between countries. The details of the EMD calculation is given in the appendix of the paper. But importantly, this satisfies the requirements I mentioned earlier, and it can be configured to weigh web sites. So this mass that's defined can be calibrated according to traffic volume, for example. The second concept we quantify is regionalization. Regionalization allows us to ask more nuanced questions, such as where providers are located and how global a provider's reach is. In some cases, it can also point us towards a causal relationship of centralization. Regionalization is partially qualitative, but we also introduced some quantifications to make it easier to interpret. We first used endemicity, a metric that was borrowed from prior work. Endemicity is calculated as the area under the horizontal line placed at the peak usage of a provider in a country. So this is telling us the deviation from absolute global consistency and usage. So in this example, if this blue mass was thinner and taller, it would be a highly endemic provider. Compare that graph to this one, which is highly popular in many regions. So in this case, the endemicity is low because it's more evenly spread out. This would be considered a global provider rather than a regional one. Measuring endemicity gives us two axes of comparisons here, which is how regional a provider is measured by endemicity and how popular a provider is, which is measured by its raw usage. So they're red and green in the graphs. So for example, a provider can be large but also regional if it's heavily used around but only in a specific region, such as Alibaba. Similarly, a small provider can be considered global if it's not used that often, but when it is, it's equally used around the world. So it's not contained to a specific region. And an example of that would be GitHub. And as expected, we have the extremely large global providers such as Cloudflare and Amazon that aren't endemic to a specific region. They're used widely around the world. The second metric we use is a pretty simple one. It's the insularity of a country or the fraction of websites for which infrastructure is served by a provider within the same country. So this tells us if dependence is concentrated within or outside a country's borders. Okay. Equipped with our metrics to understand dependence, I'll now get into our actual measurements and results. We measured infrastructure dependence across the top truck sites in a 150 countries and between three layers, hosting DNS and certificate authorities. We also conducted a longitudinal measurement two years after our initial results, and we found minimal changes in centralization scores despite a lot of churn within the crux stop sites indicating the robustness of our methods. Alright. Now we have all the setup we need, and I'm going to get into our results. Shocker. We find that hosting is not uniformly centralized. This was to be expected, but I want to highlight some interesting examples. Here's a graph of the 150 countries sorted by centralization, so most centralized to least centralized. Of note, Thailand is the most centralized country in terms of hosting with 60 of its sites being served by Cloudflare. And on the other end, while Iran is the least centralized, we see that European countries, which are the green bars, tend to be generally less centralized. And since we're in Austria, I'd like to point out that Austria ranks 140 out of the one fifty countries we studied. The next finding is that our provider classes from earlier, the regional and global providers, they provide a bit of context about what we're seeing. Large global providers do drive centralization in many countries, which was to be Cloudflare dominates in a 149 out of the one fifty countries we study, and for the remaining country, Japan, the top provider is Amazon, the underdog. But what's interesting is that the how the less centralized you are, the more you depend on a regional provider, regardless of how popular these individual providers are. So this is the same countries sorted by centralization score. The green hatched bars indicate regional providers. So this is broken down by fraction of sites and the classifications we provided earlier. And as you can see, as the centralization score goes down, the green bar gets longer, indicating a higher proportion of regional provider usage. And what this tells us is something really interesting about how we perceived smaller providers in prior literature. Take Iran, for example, the least centralized country. Most of the websites people use there aren't being served by Cloudflare or Amazon. They're being served by regional providers that wouldn't be captured by metrics that only include the top few providers. Other interesting examples include Bulgaria and Lithuania, where a single regional provider provides a significant proportion of the websites. So these are examples that highlight the importance of considering regional providers alongside the major global ones, because they often have a large say on what is available and deployed to these understudied regions. Regardless, there is a large dependence on The US, making The US also heavily insular. No surprises there. However, the cases I want to highlight are the ones in which we see cross border dependence in other regions, such as former USSR countries like Uzbekistan relying on Russian providers, former French colonies relying on French providers, and Afghani websites relying on Iranian providers. These results follow historical patterns such as with the colonies, but could also be because of shared languages causing similar websites to become popular in similar regions. So, for example, Farsi is a dominant language in Afghanistan and Iran, so you can expect that people in Afghanistan would also visit Iran websites. We see similar patterns for DNS providers, likely due to bundled services, which is why I won't go too into it too deeply. They are very similar results, but you can read more about them in the paper. The main takeaway with DNS is that provider choice can impact other layers in the ecosystem. Where we do see a shocking difference, however, is certificate authorities. Only seven CAs accounted for 98% of the websites in our data. This this is this is a shocking difference. This means that countries that are decentralized in other layers could still be heavily centralized in others. Take Austria, for example, which was, like, one forty out of one fifteen hosting, but 23 in CA centralization. And this isn't to say that this is a bad or we should be necessarily worried about it. Having fewer popular CAs means fewer entities to trust and scrutinize. In fact, this is something we'd expect to see with CAs. But from a regionalization perspective, this means most countries depend on foreign CAs, and that matters because CA infrastructure is trust infrastructure. A country can decide to stand up local hosting or local DNS, but a local CA only works globally if major browsers accept it into their trust stores. So regionalization becomes especially important here. It reveals patterns such as Let's Encrypt's popularity in Europe and, in particular, Russia's dependence on Let's Encrypt after Digicert exited following the invasion of Ukraine. Beyond service availability, the CA layer is about who gets to be trusted, which is inherently political and thus also regional. Alright. Those were a lot of results. But now that I've gone through them, I wanna slow down and highlight some key takeaways. Firstly, centralization isn't this monolithic, homogenous concept. It varies by location, layer, provider type, and it's not enough to say that the Internet is centralized because it depends on which layer you're talking about, and some dependencies could be more expected or consequential than others. Moreover, asking where these dependencies lie through regionalization results in a more nuanced evaluation of dependence. Even if dependence is decentralized across providers, it can be centralized to a specific region, which is a form of dependence that has been understudied, which is related to the idea that centralization doesn't necessarily mean that a country is insular. Sometimes dependence can cross borders, and even when it does, it's not always The US that is being dependent upon. Even though that's where a lot of dependence is, it's just one part of the broader story of dependence. Sometimes regional providers can matter on a smaller scale even if they don't appear in the global landscape. This is especially important for the standards community when measuring the reach of new standards. I want to conclude with a few next steps for this work. First, there's a lot we don't understand about causality. How does choosing one provider shape dependence on others? And when we do see high dependence on a region, what does that actually mean? Is it efficiency, lack of alternatives, market power, or something else that we don't fully understand yet? This is an important conversation that we make a first step towards, but we don't fully cover in this work. Second, I hope these metrics can help the community monitor center centralization over time. The goal is to create shared definitions and a common baseline so that conversations about centralization or just dependence as a whole are not just a collection of disconnected numbers or qualitative impressions, but grounded and measurable. Finally, I hope this work gives us a nuanced view of how far reaching infrastructure deployments actually are by identifying where dependence is located and where it's concentrated, which layers, we can better understand where to focus our efforts on to promote a fairer and more resilient web. Thank you so much, and I'm happy to take questions.

[00:27:36] Dirk Kutscher: Thank you, Rumaisa. Yeah. Let's see whether we have some questions. And, yeah, please use the echo queue. And we have Brian first.

[00:27:54] Brian Trammell: Hi, Brian Trammell. Thanks so much for this talk. I kinda wish you'd done this work about five, six years ago. The kind of wish you had done this work about five or six years ago. When I was on the Internet architecture board, we did a lot of sort of, like, thinking about this problem that actually led to the RC that you said at the beginning. This is something that we've been thinking about, like, how would you measure this? And this is the answer to that question. So thank you. Thank you very much. How would this generalize like, you're doing a little bit of sort of layer layer hopping around here. Right? You're like, here are the various points in the in in the network and the nodes. How would this generalize beyond the layers that you've chosen? So

[00:28:35] Rumaisa Habib: the framework that we mentioned and the quantifications was, like, purposely more general. So Mhmm. The metric, the EMD metric, it could be any doom distributions. We're just preparing against a fully distribute a fully uniform one, and it can be weighted according to whatever the needs of that layer is.

[00:28:53] Brian Trammell: Right. Any any plans to do to to broaden this?

[00:28:57] Rumaisa Habib: To do broad To

[00:28:58] Brian Trammell: to broaden this, like, a little bit, like, so up to the application layer a little bit and then down. Like, so Yeah. For example, like, there's a bunch of other things that I wanna talk about that I think I'll take to the IRSG dinner because there's a queue. There's a bunch of other things you could you could say about, like, how we got there on the distribution of, like, Cloudflare being where it is and how we got there on the distribution that would lead to other sort of questions that you could ask the dataset.

[00:29:22] Rumaisa Habib: So I would definitely be interested in measuring

[00:29:24] Elliot Lear: Thank

[00:29:24] Jonathan Hoyland: you very much.

[00:29:25] Rumaisa Habib: Thank you.

[00:29:27] Speaker 5: Hi. At Netflix. I think Brian touched on some of the things that I was asking. But more specifically, I was curious, have you considered adding more like an AS topology network level path view and dependencies and and upstream providers for for country level dependencies?

[00:29:43] Rumaisa Habib: Yeah. There are definitely different layers that we should also study. And there are other papers that also look at these. There's a paper on DOM dependencies, third party resources on the web. So, like, yeah, it's definitely applicable to that as well.

[00:29:57] Marwan Fayed: Alright. Thank you.

[00:29:58] Rumaisa Habib: Thank you.

[00:30:00] Elliot Lear: Hi. Elliot here. Ramessa, that was delightful. Really, really well done. I have a question about the CAs that which is your your findings were disturbing, actually. Right? I mean, it it it's seven players and what happens if there's a common fault between them, for instance, which it's possible. My question is, in terms of Let's Encrypt, did you measure its penetration in particular versus so of that 98%, how much of that is is Let's Encrypt?

[00:30:35] Rumaisa Habib: I don't have an exact number for you, but I think it was either the I think it was the second most popular or the first most popular and had a greater usage in Europe in particular.

[00:30:46] Elliot Lear: Yeah. So a comment here for the community to think about, right, is Let's Encrypt was essentially founded out at Us geeks.

[00:30:54] Rumaisa Habib: Mhmm.

[00:30:54] Elliot Lear: And it's been wildly successful, maybe too successful. Your point about the balance of trust between the endpoint and regionalization goals. Right? Could Let's Encrypt be restricted, for instance, by sanctions or other things? I think we have ourselves a

[00:31:15] Rumaisa Habib: problem. Yeah.

[00:31:17] Elliot Lear: That you've that you've identified. Right? I mean, that that's that concentration is is Yeah. It may be a little too concentrated.

[00:31:25] Rumaisa Habib: Thank you for the comments.

[00:31:27] Marwan Fayed: Hi, Romesa. So one is, you know, how I feel. I think this is really fabulous because it helps shift the conversation. So I'm gonna echo comments you've heard before, maybe a bit of a soapbox, which is I think shifting the conversation away from the c words being something bad to something that that maybe deserves attention in how to manage is important. For me, if I think of the past, I can't find any evidence of any large infrastructure built by humankind that moves things around, that manages to scale successfully without points of concentration and centralization.

[00:32:02] Rumaisa Habib: Yeah.

[00:32:03] Marwan Fayed: To me, that's just an invariant, and then we have to shift the conversation towards moving about how to manage it, how to monitor it, how to safeguard it, these types of things.

[00:32:11] Dirk Kutscher: So thank you.

[00:32:12] Rumaisa Habib: Yeah. I fully agree. And this quote is, like, pretty important in that discussion. It's like, it's not enough to say if something is centralized or not. It's about the outcomes that you see. So I agree with you there.

[00:32:26] Ignacio Castro: Thank you. Ignacio Castro from Queen Mary University of London. Thank you very much for the cool presentation, and I think this question is a bit of a follow-up on the one from Marwan. So is centralization inexorable, or it can be affected? And if so, how? And if there is a how, is that technical, political, or social? And I realize that probably there is not a very direct answer to this, but your thoughts would be more than welcome. Thank you.

[00:32:52] Rumaisa Habib: I I think it's both definitely. Like, it is technical. It's political. Like, how do you manage to make infrastructures to speak to each other? Like Marwan was saying, like, to be successful, you need, like, successful ways to communicate with each other as well. So it's definitely both.

[00:33:09] Ignacio Castro: Anything specifically that this community could help with?

[00:33:12] Rumaisa Habib: Sorry. I didn't catch that.

[00:33:13] Ignacio Castro: Anything specifically that this community could help with?

[00:33:17] Rumaisa Habib: I think standardizing what we use to measure it and so we can monitor it over time is would be very helpful.

[00:33:28] Sam: Hello. I'm Sam from University of Twente. So I was considering cases where, like, big providers such as Cloudflare, it puts its server in a local data center in a country. So how did you analyze such cases? Because in terms of dependency, it looks like, the server, the website is hosted within the country. So how often did you find such cases, or did you consider analyzing those cases?

[00:33:56] Rumaisa Habib: Yes. I didn't have the chance to talk about it in this talk, but the paper does do IP geolocation to understand that as well. So you should definitely check that out.

[00:34:05] Sam: Alright. Thanks. Thank you.

[00:34:08] Jonathan Hoyland: Jonathan Hoyland. Is there a meaningful way to combine the scores you got for different layers? Like, if I'm super dependent on country a at layer one and super dependent on country b at layer two, is that different from being dependent on the same country at both layers, or is that the same?

[00:34:33] Rumaisa Habib: I think it's very different. Like, the context of dependence on the CA layer, for example, and the hosting layer, it it's a completely different ecosystem. And I think we would it would be disingenuous to combine them and say this is because it's not the full picture. Like, you like, the example of Austria that I gave, like, it's pretty decentralized in hosting, but then when it comes to certificates, it's not. And if you combine that, we'd probably find that somewhere in the middle, and that's not helpful to continue the conversation.

[00:35:03] Dirk Kutscher: Thank you.

[00:35:04] Rumaisa Habib: Thank you.

[00:35:06] Dirk Kutscher: Okay. No further questions. Then let's thank Rumaisa again. Okay, thanks again. Great presentation, big paper.

[00:35:44] Diwen Xue: Congratulations.

[00:36:09] Dirk Kutscher: Okay. So we continue with the next A and R P talk, and this will be by Diwen Xue on deep packet inspection devices by their ambiguities. And Diwen is a recently appointed assistant professor in the department of information engineering at the Chinese University of Hong Kong, CUHK. Before joining CUHK, he was a research fellow at the University of Michigan Ann Arbor, and his research sits at the intersection of security and privacy and systems and networks. And you have the rest of his bio in the presentation. Thank you.

[00:36:57] Diwen Xue: Thank you.

[00:36:58] Dirk Kutscher: I bring your presentation up. Yep. Just a moment. Okay. Take it away.

[00:37:24] Diwen Xue: Yeah. Can can you guys hear me okay? Cool. Yeah. Thanks for coming. I'm Diwen. Good afternoon. I'm from the Chinese University of Hong Kong. And today, I'm going to be talking about our recent work measuring deep packet inspection devices. So this work was done over a year ago when I was a PhD student at the University of Michigan. Alright. Before I start, I want to know just one thing. Fingerprinting has been overloaded with different meanings these days. And in this work, means it reads it it it reads like, you know, differentiation and telling things apart more than tracking. Just one one note. So yeah. In recent years, users' Internet traffic has seen more and more interference from firewalls and middle boxes along the network path. We are talking about censorship, filtering, and throttling becoming more common. And it's really not just just authoritarian countries. You know, around the world, are seeing more and more restrictions in the form of geo blocking, server side blocking due to sanctions, and new legal frameworks that institutionalize these practices. This trend we've been seeing is largely in or at least in part driven by the growing availability of deep packet inspection devices or DPI's, which are designed to monitor, inspect, and filter network traffic. So the the commoditization of DPI technology has made these devices cheaper, more capable, and essentially, we've seen the reach of any network operator. So, of course, they serve many legitimate purposes like intrusion detection, but that same capability also made them the primary enabler of many network interferences that we call censorship, like filtering, throttling, broadly defined. The DPIs are fundamentally tall use technologies. Right? And as with all dual use technologies, it's important to maintain maintain visibility and transparency around how they are deployed and how they are being disseminated. So let me start with a concrete example. Back in 2021, some Russian users began to notice a slowdown to their Twitter connection to their Internet connection. And later on, we found that it was a intentional large scale throttling against certain social media services, like including Twitter, to in to discourage their use. So back then, we tried to, you know, measure this incident. You know, we started with the effects of the interferences, like, you know, characterized the the effect of throttling as being implemented. We for example, we record and then replay some network activities, and that allow us to distinguish between intentional interference, like throttling, from the background noise in the network. So for those of you who are interested, I'll refer you to our previous work measuring throttling. But those but the previous slides, those observations so far told us that there was a DPI device, right, somewhere along the network path that was interfering with our connection. But how can we know more about these DPI devices themselves? So this is a research question that remains, I I would say, is very understudied even today, because there has been many work measuring the effects of the interferences, like what they be what what are being censored, what are being blocked, what are being throttled, like what we did in the previous slide. But very little work focused on the devices responsible. Who manufactured them, how they are being deployed, and how they are how widely they are being disseminated. And this this gap is really here for good reasons, actually. You know, first of all, the DPI devices are network intermediaries. Right? They sit in the middle of the network and then makes them almost invisible to standard measurement techniques that we usually add our tool sets, a ZMAP or NMAP, which are designed to send endpoints. Most DPI's deployments, they do not expose public IPs that you can directly pin, and their presence is, to some extent, transparent when they are not being triggered. So the first challenge is that standard measurement techniques out there are not applicable, directly applicable to measure on past DPI devices. But more interesting but more interestingly, like, the DPI vendors themselves have a strong incentive to stay hidden. Right? Like, this often intentionally try to obscure their devices to avoid explicit identification. But this is especially true after a series of lawsuits in recent years that has exposed the role of certain, I would say, like, certain Western made DPI products in facilitating censorship and surveillance in certain regions of the world. Like, companies like Fortigate, Blue Coat do not wish their names to be associated with, for example, doing throttling or censorship mostly because of the controversy around doing these interferences. So this incentive to stay hidden really shows up clearly now in our data. Right? Like, before this work, a common way to identify DPI devices was through the blog pages they inject, like, which sometimes contain vendor identifying information like a a logo or key phrase or product name. The previous work, for example, our own filter map from 2020 used these blog pages to attribute censorship events to the devices responsible. Sorry. But but ever since 2020, this this signal has been disappearing. This figure here shows is our longitudinal measurements, which shows the share of the the share of censorship that comes with a explicit blog page has been reduced by over 85% since 2020. Because overall censorship is actually increasing, but just its place in block pages, especially the ones that contains vendor logo and, you know, a a product name has been replaced by more generic methods, like a generic TCP resets, which now accounts for 95% of all the blockings we see around the world. So that brings us to the to the research problem of this work. You know, when the DPI devices are becoming more and more generic, when their actions are becoming more and more generic, how can we still tell them apart tell tell the DPI devices apart? So before I talk about our method, I want to go back to the throttling example because that is also where the idea of this from this this work came from. One of our main question back then was that whose devices were doing the throttling. So to give you a little bit more context to understand this, Russia for a long time followed a so called decentralized censorship model where it's very actually very similar to many Western countries where you have a centralized block list, but the enforcement of the censorship is left to each individual ISPs using their own commercial DPI devices. So but around 2021, there was a a rumor going on that Russia was moving towards a more centralized model where the censorship censorship will not be carried out by individual ISPs anymore, but rather by a new set of new DPI devices that are centrally deployed and centrally controlled by the government. So our research our one of our research question back then was that who was doing whose devices were doing this were implementing this throttling. So at first, we tried a bunch of things with with with little success. Right? Like, the throttling itself, like blocking or other type of censorship, carries very little information about the device enforcing it because it's a very generic interference. Then we found something useful,

[00:45:49] Elliot Lear: you

[00:45:49] Diwen Xue: know, to some extent by accident by accident. We were testing how the throttling device would handle IP fragmentation reassembly, and we noticed that the device had a very specific limit when it comes to IP fragmentation reassembly buffer. So it would only more specifically, it only would sorry. It would only accept fragments up up to 45 fragments for a single IP packet. So if you split split the throttling trigger into 45 fragments, you can see the throttling. If you split the trigger into 46 fragments, you no longer see the throttling. So lucky for us, this specific 45 fragments, this this fingerprint, was quite unusual. We checked other mainstream firewall products out there, Cisco, Juniper, and their limits are, like, 24, 32, or 64. Right? And at the same time, we also found that across the entire country of Russia, across hundreds of privately owned ISP, we found this you know, the devices doing all share this same unusual 45 fragments as a limit. So in that work, we conclude, like, know, this high level of uniformity did not really fit into our previous understanding of that, you know, decentralized censorship model and instead points to something that is more, you know, is a single coordinate deployment. So the result is that from that, the fingerprint this fingerprint over IP fragmentation reassembly buffer eventually allowed us back in 2022 to map out the deployment of this new national firewall for Russia that no one had described before back then. So this is the idea. So this overall, this is idea that we want to generalize in this work. You know? Not not not not one quirk found by that only worked in Russia, but a way to do this at scale to for any DPI device around the world. So following this motivating example, we built DMAP, which is a measurement framework specifically designed for measuring the pet DPI devices. So because DPI's are network intermediaries, measuring DPI's is fundamentally different from measuring a end endpoint. For example, for typical Internet measurements, you send specially crafted packets called probes directly to your measurement targets, like a web server, and these probes would trigger some replies, and then you analyze the replies to learn something about the web server. So for dMAb, we still need to point our probes into some general directions, like a web server, but the real target of these measurements is not the web server per se, but rather some middle boxes rather some DPI devices along the network path. So the explicit reply to our pros probes is not useful to us because it is generated by the web server, not by the DPI itself. And in these measurements, the only thing that is useful usually come usually comes down to the observation of some side effects, like in our previous example, throttling at the presence or absence of throttling, which is just a binary signal. Know, throttling triggered versus throttling not triggered. And what dMAb does is to use this binary signal as a side channel to gradually to infer one bit of information each time to gradually build a fingerprint to help us differentiate and cluster DPI devices. So that is a high level idea behind dMAb. So, yeah, more specifically, this is the framework looks like. And the more important the most important part of this is what do we use as fingerprints to differentiate DPI devices? So to build a fingerprint for anything, you need a source of entropy. Right? Something that is different from one implement one DPI implementation to another. But in our previous example, the size of the fragmentation reassembly buffer was one of was one such source of entropy. The d map tried to generalize over this. We call the source of entropy ambiguities. So essentially, each ambiguity is a sequence of packets whose handling is not well defined, and it's likely being handled differently across different DPI's. And that those differences, those entropy are what allow us to differentiate and fingerprint each unique implementation. So let's let yeah. Let's look at these ambiguities for for a little bit. In an ideal world, we would like to all the devices on the Internet probably should handle traffic exactly the same way. But in the real world, as I'm sure all of you know, there are ambiguities and corner cases for several reasons. Right? Right? For example, protocol RFCs cannot specify every possible behaviors, and that leaves that forces DPI vendors to make make some of their own design decisions, which can be really messy. And even when there is a specified behavior, the ambiguities in natural languages, as we will see later, leave room for different interpretations. So the same standard can be implemented differently. So take the TCP take the DPI's handling of TCP timestamp as an example. Right? We found even among the most popular open source DPI's out there, Zig, Suricata, what was it, Snort, and NDPI. They don't even agree on the basic questions, like, whether they should process, whether they should care about this option at all. If so, how to parse it? What is a valid time? And what to do if the time is not is not valid? These implementation discrepancies are a well known issue in software engineering, but here, we leverage them for measurement. Yep. I won't get into the details here, but, basically, we built a ground truth test bed with dozens of commercial and open source deep DPI devices with labels. So then we did differential fuzzing over this test bed by enumerating all the fields and aspects of connections that we can mutate, and then look for mutations that are handled differently across our test bed. So we found thousands of these candidate ambiguities with each one providing some amount of entropy across different DPI's in our test bed. So, yeah, one example here, we this is one of top ranked probes, one of the ambiguities with highest entropy. So we set the in this, we set the sequence number of a repressed packet to a negative value relative to initial sequence number, but also pass the beginning of the payload. Right? So the so this probe cost different handling across different DPI devices. So we look into the open source ones that we look look into their source code. We found all of them would discard packets that have invalid sequence number, but their definition of invalid is different. Some like Zeek do Zeek do simple checks that compares sequence number to the initial sequence number, while others do a more nuanced chat like compare like snorts that compares the beginning of this of the payload to the upper bound of the receiver window and also the the end of the payload to the lower bound of the receiver receive window. So in that sense that even if the with a negative sequence number, the part of the payload that is in window, it's still valid for reassembly. So in this case, whether the DPI accepts partially out of window segments provides one bit of information to help us differentiate implementations like z from, for example, snort. This is another example. This one is simple. We just said send a certain the timestamp of the set certain patches to be lower than the the previous pat packet. And this about the this is about a mechanism called pulse, protection against wrapped sequence number. So some DPI's we tested in our test bed correctly implement this mechanism. Some like c trans NDPI do not implement this at all. And interestingly, for snort, it does implement this mechanism but in a wrong way. Right? So so ROC requires to, for example, invalidate the check when the connection has been idle for too long, but apparently, Snort developers interpreted invalidate to mean that invalidated packet, not the check, which means that they are implementing the exact opposite behavior of what RFC requires. So what does this mean? It means for all for for the for these three open source APIs, none of them is compliant with RFC. Like, they are all wrong, but they are wrong in in different ways, and they are wrong in different and fingerprintable ways. So after we collect these ambiguities, right, by doing differential testing on our test bed, we can start using them. We combine the ambiguity with a known interference trigger. Right? Going back to the throttling example again, this means collocating the ambiguity with a known trigger, for example, the SNI or domain name that triggers throttling. Remember, the only observable signal that come from directly from the DPI is just this observation of some it's it's just the binary side effects. Right? The presence or absence of interference. And now by observing this binary side side effect, we can infer how the ambiguity is being resolved by the DPI. So in other words, for each ambiguity, we have a binary outcome, you know, interference triggered, not triggered, ones and zeros. And finally, we combine these ambiguities we combine these binary outcomes across different ambiguities into a fingerprint. So if we visualize this measurement, whole measurement process, imagine these are all these are all the DPI devices in on the Internet. The most ideal probe that contains an ambiguity that these devices strongly disagree whether to process it or ignore it. Right? So half so, like, the two examples I just showed you, so half we observe a one and half we observe a zero. So essentially, each probe would each ambiguity would divide all the DPI devices on the Internet into two partitions and depending depending on how that ambiguity is being resolved. And with more probes, these partitions become smaller and more homogenous. So eventually, the fingerprint is long enough and unique enough for to to identify each unique implementation. So now now let's look at how this measurement works in practice. Right? Start with something small. We bought a a and, like, we we we got a a bunch of DPI products off the shelf from different vendors, and we we test them with 40 ambiguities. So the top top ranked 40 ambiguities. So each DPI product has a 40 bit fingerprint. You can clearly see that DPI is made by the same vendor cluster closer together, you know, which tells us that their implementations are similar, maybe sharing parts of their code bases. Now if we increase our scale, these are random these are random samples of 40,000 firewalls that we collected from we found on the Internet using sensor plan as public measurement data. We cluster them based on their d map fingerprints and then color by their geolocation. So you can see that most clusters indicate original deployments. Like, we found multiple clusters corresponds to national firewalls or local vendors, but we also found clusters that we also found some deployments that are more global, like a specific FortiDate product been deployed in over 30 countries around the world. Yet we also asked how often do these fingerprints actually change over time. And for this, we look at one two of the most popular open source APIs out there, Zeek and Suricata. We and we track their implementation over the last five years over more than 80 software releases. We found for Zeek, the the fingerprints was essentially identical throughout its entire period of five years. So across, I don't know, I forgot, like, six or seven major versions. And for Suricata, we only noted there were only two noticeable changes over the over the 80 releases. So which means that these fingerprints might not change frequently enough as to require constant remeasurement. So, of course, this method is not perfect. It has some limitations. Right? And and one of them is that we are doing all of these things not in vacuum, but in a very noisy environment. Right? So our method depends heavily on DPI devices being triggered reliably. Right? Because we are using that as a binary feedback channel. But in some environments, DPI's do not trigger reliably, and that adds a lot of noise to the measured fingerprint. For example, we found DPI's in we found targets in China, Turkey, and I believe Cuba had much more much higher variation within their clusters than other tar than targets in other countries. So we found that censorship in these regions do not trigger consistently even if you're using the same a trigger domain and just it's a in in the exact package sequence. So this probabilistic triggering caused a lot of so called, like, beat flipping in the measured fingerprint. A simple way to reduce this noise is to do repeated measurements and then aggregate the outcomes by prioritizing ones over zeros, you know, under the assuming that most DPI's fail open rather than fail close. And this indeed we do makes our fingerprints much cleaner at the expense of more measurements, of course. But a more significant but more significant limitation comes from the basic nature of remote measurements. Right? Whenever we send a probe and observe an apparent censorship outcome, we have to ask whether this is was because the ambiguity has no effect on the DPI device or whether there were some entities along the network path that discard or normalize our packets in cells. So we are essentially fingerprinting the normalizer, not the DPI. Or maybe the end host is simply not responding. For some of these for for some of these cases, we can protect ourselves by carefully design, test, and controls. Like, so for each bit of the fingerprint, we are not running one set of measurements but four. So using a trigger domain, a benign domain, and then with and without the ambiguity. And so we look at the pairings of those four sets of responses. Only some pairings have a valid interpretation. You know, other combinations will tell us if something went wrong, like maybe there are multiple DPI's along the network path or the end host has interfered with our measurement. And in those cases, we threw the measurements results. And, of course, we can also add TTL limited measurements. So to to summarize, we absolutely need to have a better understanding of this network intermediaries, this DPI devices on our network. And the challenge is becoming harder because these devices are becoming more and more generic. They are sending fewer and fewer self identifying blog pages, which means the previous methods that worked well sorry, which means the previous methods would not work as well anymore. So what we have done in this work is really just a first step, you know, to be able to remotely differentiate and cluster DPI devices based on their implementations. So the main insight of this work is is is to shift the focus, right, from what DPI's write to the wire to focus on instead how they read the traffic because what they write is limited, and it's becoming more and more gen generic as we speak. But how they read traffic and parse and reassemble is this large space of full of ambiguities and entropy. So we hope this work brings more measurement efforts to to bring eventually bring more transparency and understanding into how these devices are deployed, how these these DPI devices are deployed, and how they are disseminated on our Internet. So thank you for listening, and I'm happy to take any questions.

[01:04:03] Dirk Kutscher: Thank you, Duen.

[01:04:12] Speaker 10: Hey. Hi. Hi, Duen. Super interesting talk. Some fifteen or twenty years ago, I was working on DPI things, and I kinda left it when I thought we were gonna encrypt everything and there wouldn't be anything interesting there and focused on just things we should be reading like headers. The two questions I have for you is the first is, what's the value to these companies like FortiGate and Fortinet or Blue Coat when unless they're man in the middling the certificate. If they can't decrypt the payload, what's the real value to a customer of this? Are they just after the server name identification and making the decisions there? Because I get that that's not very deep, but it is deep. So that's the first question. And then the second is, I went down some rabbit hole in my new job that I went to last year and a half, and I found a Forta something product interfering with I p v six traffic in a completely nondeterministic way. Like, I was wasting hours.

[01:05:12] Diwen Xue: Interfered with what?

[01:05:13] Speaker 10: Maybe I'm I'm going too fast. Let's let's just go the the first question. Yeah. What's what what are they after that's not that's different than the SNI, or is that all they're after?

[01:05:22] Diwen Xue: So let let me rephrase your question to make sure I understand. You're asking why they were doing its Placid blog pages in the first place that would review their identities?

[01:05:30] Speaker 10: What's the value of buying a a deep packet inspection product if you're not man on the middleing and and decrypting the pay payload? What what what is anyone getting out of this product?

[01:05:42] Diwen Xue: I'm I'm I'm sure sorry. You you I don't quite get your what you are asking. So you are so what is the value of buying DPI products if you can What

[01:05:50] Speaker 10: are they discovering in the payload other than SNI if it's encrypted? What what's what's what's the point of me going out today and buying a deep packet inspection product?

[01:06:02] Diwen Xue: So both of these filterings and, you know, throttling all this, you know, interference are triggered not on the metadata, like the SNIs and all the things that they they don't have to decrypt the payload. If I'm not sure I understand your question, honestly.

[01:06:18] Marwan Fayed: I so, Buen, is I I can help bootstrap this. Example, I run Apple Pie Corp, and I don't want people to know anything about Blueberry Pie Incorporated. It's it's really that simple. It's it's the SNI that people are after in a lot of cases.

[01:06:29] Speaker 10: Okay. That's useful. And then maybe the other is just an anecdote that I I wanna share if anyone can help me. I I because these are hidden boxes often, when they interfere with the traffic, it is so difficult to find out what they're doing. So I found, for instance, that they occasionally interfere with running ipv6 tests to test ipv6.com. And then you're saying, like, one out of a 100 times, one out of 20 times it happens, and I'm at a loss. So I was wondering if you had any insight as to what is it doing. Is it is it confused about I p e six? Does it sometimes get false positives on encrypted payloads that it should never be trying to test anyway? Thanks.

[01:07:07] Diwen Xue: Okay. Yeah. Thanks.

[01:07:09] Dirk Kutscher: Quick note. So the acoustics here are really horrible. So you have to speak a little bit slower than normal. Otherwise, we will not not understand you.

[01:07:18] Diwen Xue: Yeah. Yeah. Can you speak to a little bit louder, please? I also have a hearing problem. Sorry.

[01:07:23] Sam: Hello. I'm Sam from University of Twitter. There are security devices which which acts on behalf of devices behind them, which means they allow normal TCP connections. After that, they reply with Windows zero, which means the purpose is to exhaust the source or, let's say, attacker or measurement method methodologies. So did you consider looking at the cases of Windows zero? Would that help with improving your detection methodology?

[01:08:02] Diwen Xue: I'm sorry. I really can't hear your question well. Maybe we can Yeah. Yes, sir. Discuss later on. Okay. I'm so sorry.

[01:08:14] Stuart Cheshire: I'm Stuart Cheshire from Apple. I'll try to speak slowly and clearly. Can tell the acoustics are bad. Thank you for presenting this work. Several times, you talked about throttling traffic, and I'm curious about that because if we don't think about it very much, it's easy to have an intuition that if you squeeze a garden hose, then the water comes out slower. But packets in the network are not like that. If you have a piece of hardware, you can't stop the packets arriving over the wire, the fiber, radio waves through the air. The packets will come in whether you like it or not. And I think your choices are you can forward the packet. I guess you could delay the packet and then forward it. You could modify the packet and forward it, say, the ECN bit in the header, or you could discard the packet. Do you have any insight when you are talking about Twitter being throttled? What was the mechanism they're using to produce that throttling?

[01:09:27] Diwen Xue: Oh, I think so that was a different word, but I remember back then it was like a token bucket implementation that would discard once you reach above their limit, all the other path here within the next one second would be discarded. Discard.

[01:09:46] Stuart Cheshire: Thank you.

[01:09:52] Dirk Kutscher: Marwan, do you have a question? I had to hand up in the fax.

[01:09:56] Speaker 10: I'm sorry.

[01:09:56] Elliot Lear: Okay. Okay.

[01:09:59] Dirk Kutscher: Any other question? Okay. Then, let's thank Diwen once again. Okay. Thank you. Thank you. Good luck. Great. Thank you. Okay. Just a moment. Okay. That was great. Let me give you quick update of what happened earlier today in a site meeting. So if you remember, we have been discussing what to do in the IRTF about networking for AI for a bit of time. And so today, we we had a side meeting here and discussing with scoping ideas what could be a good topic. And so the background of the discussion today was, so as you are surely aware, there are many things going on. So

[01:11:37] Elliot Lear: in

[01:11:37] Dirk Kutscher: the ITF inside meetings around the general topic of AI, agent communication, and so on. So for the IETF, so we see there are some attempts to start some, let's say, engineering work soon on things like agent gateways or the agent portal above earlier this morning. So this is typically short term engineering efforts. And so what was discussed today is there room, but is there also a utility to have a longer term research effort in this larger field? So not only agentic AI, but also AI infrastructure, so network systems for AI. And let me skip a little bit to these these group summaries. So in general, it's quite a bit of enthusiasm. So today, but also in the discussions earlier. If you look at what's currently going on in like the top tier networking research venues, So there are a lot of contributions on improving infrastructure for different aspects, AI training, inference, distributed AI systems, and so on. And so the question is, what of this work is relevant? So in in the ITF, IRTF context. And so what could be a useful way to work with these communities? And so we we think so that was a bit the the feeling in the room, at least, that, well, there's quite a bit of, you know, Internet technology research going on there as well. So, like, using protocols, using techniques like packet spraying, receiver driven protocols, and so on. And that are currently used in experiments research. And there's quite a bit of confidence that it's it's interesting to evaluate these things, make recommendations for future work. And then there's also the question, you know, so what could be like topics that are maybe don't get enough attention perhaps or have like more potential for for future work? So, you know, more principled decentralized system in this space, more principled architectures for agent communication. There are some research projects that have been kicked off, like Internet of Agents and similar things. And then also, as things are moving a bit out of a data center or as also data center technologies evolve, is there room for more internet native support for AI systems? And so we discussed these things a little bit. You will find the details here on the slide. Let me just maybe skip the details for now. So potential activities that we discussed were, first of all, maybe also investing a bit more effort in creating experimental facilities. So tools, simulators, test beds perhaps, that could enable better experiments and also better evaluation of different systems. And then in the same context, also maybe making data sets available. So other communities do this quite successfully, so to enable reproducible experiments. And we thought that could be potentially interesting for this topic here as well. There was a bit more interest in not only having a group that invites interesting talks, but really tries to you know, push the needle a little bit and also lead lead research in this field. So last, IETF, for example, we had an invited talk by the Mooncake company. It's Kimi K3. And they also gave a talk this week in ICCRG on their KV cash centric system. So people think that there's good potential to kind of think this further, maybe also in a more distributed way. And then also for like large language model inference, there are like mixture of expert architectures that are constantly evolving, which could also be a good topic for for our work here. And then we also now see new protocols that actually address security, which was maybe previously left out because not deemed important in data center environments. So we had an talk at the last ITF about a corresponding protocol. And yeah, the other big topic that was discussed was developing principled approaches on agentic communication. So not so much focused on kind of reusing existing technologies necessarily, also or fitting into existing deployments or business models, but really a bit more fundamentally. So how could a secure, trustworthy, scalable Internet of Agents be conceived? And, yeah, overall, there was quite a bit of enthusiasm. There are still some open questions regarding focusing or broadening the scope. So what's the right right scope here? But I had quite a positive impression of of that meeting. And I just wanted to open it up now for any further questions or comments if people have ideas on this topic, so I have any interest. So the main purpose was to inform you what's currently going on.

[01:18:21] Elliot Lear: Elliot. Dirk, thanks for this. And I'm looking at that Internet of agents, and I'm thinking, oh, I wonder if that's really a a federation, you know, topic or something along those lines or groups of federations or the the the scaling trust aspect there is really, really a meaty research topic. The one thing I just wanted to mention with my Internet with with my independent submissions editor hat on is that I have a a good number of AI related drafts. Some of them are around shared context. So maybe those go to the IETF if they're engineer y, or maybe maybe you still want them here if they're a little more research y. But I do have some that are focused more on transport of secured transport of information. Maybe those go here or maybe those go there. There's a a theme forming here for those people who don't know notice it, which is that the independent submissions editor is not disposed to be the starting point for this work. Right. So if there's opportunity for to find a home for it, I think it's a good organizing principle that we do so, and I encourage the formation of a of a broader research a a broad research I won't say broader. That looks pretty broad. But but broad research and and for the IET, of course, to appropriately organize. Okay. Thank you, Elliot.

[01:19:52] Dirk Kutscher: Okay. Yeah. That's it. Thank you very much for coming and see you again next time.

[01:20:43] Rumaisa Habib: I know to see a drumbeat.