Skip to main content

According to the Best Available Data

The CAIDA blog — commentary and analysis of Internet measurement, infrastructure, policy, and economics, published since 2006. Formerly hosted at blog.caida.org. See About the CAIDA Blog · Categories · All entries by month · RSS feed.

CAIDA's Annual Report for 2011

[Executive Summary from our annual report for 2011.]

This annual report covers CAIDA's activities in 2011, summarizing highlights from our research, infrastructure, data-sharing and outreach activities. Our current research projects span topology, routing, traffic, economics, future Internet architectures, and policy. Our infrastructure activities continue to support measurement-based studies of the Internet's core infrastructure, with focus on the health and integrity of the global Internet's topology, routing, addressing, and naming systems. We are also dedicating resources to support the infrastructure measurement and data sharing interests and needs of two U.S. federal agency programs: the National Science Foundation's International Research Network Connections (IRNC) program, and the Department of Homeland Security's Protected Repository of Data on Internet CyberThreats (PREDICT) data-sharing project.

We continue to expand our Internet active measurement platform Ark in scale and functionality, and use this platform to collect and share the largest Internet topology data sets (IPv4 and IPv6) available to academic researchers, and share many aggregated annotated derivative data sets publicly. Our topology measurement platform supports IPv6 -- by the end of 2011, 28 of our 57 Ark hosting sites provided IPv6 connectivity and topology measurements. We have dramatically improved existing techniques for IP address alias resolution for large Internet graphs; we submitted a paper describing and evaluating the performance of our algorithms in late 2011, hopefully for publication in 2012. (Preliminary technical report available on the web site now, see Topology section of the report.) Using these new techniques, we collected, analyzed, processed and released two Internet Topology Data Kit (ITDK) Datasets, reflecting measurements taken in April and October 2011. Each 2011 ITDK includes two related router-level topologies, router-to-AS assignments; geographic location of each router; and DNS lookups of all observed IP addresses. We are still working on improving and validating our AS relationship inference algorithm so that we can add additional annotations to future ITDKs.

On the theoretical side of topology research, we continued investigation of the geometric model we developed last year to study the structure and function of complex networks. This model assumes that hyperbolic geometry underlies many complex networks, which if true provides a natural explanation for the heterogeneous degree distributions and strong clustering that characterize so many complex networks, i.e., they are simple reflections of the negative curvature and metric property of the underlying hyperbolic geometry. We also showed that not only popularity but also similarity acts as a strong force in shaping complex network structure and dynamics. We developed a framework where new connections, instead of preferring popular nodes, optimize certain trade-offs between popularity and similarity. The optimization framework more accurately describes large-scale Internet evolution (new links) than previous models, e.g., preferential attachment. The mathematically inclined will appreciate our related recent investigation of random bipartite networks using a hidden variable formalism that facilitates study of the structure and function of complex networks, as well as inference of individual characteristics, attributes, and annotations of nodes in real bipartite networks. Particular applications of interest are network geometry and navigability.

We gained momentum on our economics and policy research agenda, focused primarily on explanatory and predictive modeling of the economics of transit and peering interconnections in the Internet. Two historical developments contribute to a persistent disconnect between economic models and actual operational practices on the Internet. First, the Internet became too complex - in traffic dynamics, topology, and economics - for currently available analytical tools to allow realistic modeling. Second, the data needed to parameterize more realistic models is simply not available. The problem is fundamental, and familiar: simple models are not valid, and complex models cannot be validated. We are making progress in both dimensions: creating more powerful, empirically parameterized computational tools, and enabling broader validation than previously possible. We also held the second interdisciplinary Workshop on Internet Economics (WIE) in December, connecting academic researchers, commercial Internet facilities and service providers, theorists, policy makers, and pundits of Internet economics to frame an Internet economics research agenda, and more specifically to improve the realism, utility, and predictive power of economic models of Internet topology and dynamics.

In the first months of 2011, Internet communications were disrupted in several North African countries in response to civilian protests and threats of civil war. We analyzed episodes of these disruptions in two countries: Egypt and Libya. Using both control plane and data plane data sets in combination allowed us to narrow down which forms of Internet access disruption were implemented in a given region over time. Among other insights, we detected what we believe were Libya's attempts to test firewall-based blocking before they executed more aggressive BGP-based disconnection. Our methodology could be used, and automated, to detect outages or similar macroscopically disruptive events in other geographic or topological regions.

We are applying our theoretical, empirical, and practical understandings of the Internet's evolution to engage in the NSF's exciting Future Internet Architecture (FIA) Research program. In 2011 we participated in the Named Data Networking project, a 12-university collaboration funded by the FIA program to explore a generalization of the Internet architecture that allows naming more than just communication endpoints, i.e, the source and destination IP address, but also data (content) itself. This approach shifts the focus from where -- addresses and hosts in today's Internet -- to what -- the content that users and applications care about. By naming data instead of locations, the new architecture transforms data into a first-class entity while addressing the known technical challenges of the today Internet: routing scalability, network security, content protection and privacy. In 2011 we investigated combinations of name-space structure and network topology that optimize the efficiency of NDN algorithms and participated in NDN testbed development and evaluation.

Finally, as always, we engaged in a variety of tool development, data-sharing, and outreach activities, including web sites, peer-reviewed papers, technical reports, presentations, blogging, animations, and (six) workshops. Details of our activities are below. CAIDA's program plan for 2010-2013 is available at http://www.caida.org/home/about/progplan/progplan2010/. Please do not hesitate to send comments or questions to info at caida dot org.

Full annual report:
http://www.caida.org/home/about/annualreports/2011/

Program plan for 2010-2013:
http://www.caida.org/home/about/progplan/progplan2010/

IPv6: What could be (but isn’t yet)

With IPv6 Launch approaching, there is increasing interest in measuring the readiness of the IPv6 infrastructure. A major concern, particularly for networks that source or sink content, is the performance that is achievable over IPv6, and how it compares to the performance over IPv4. A recent study by Nikkah et al. argues that data plane performance, as measured by web page download times, is largely comparable in IPv4 and IPv6, as long as the AS-level paths in IPv4 and IPv6 are identical.  We have confirmed these findings with our own measurements covering 593 dual-stack ASes: we found that 79% of paths had IPv6 performance within 10% of IPv4 (or IPv6 had better performance) if the forward AS-level path was the same in both protocols, while only 63% of paths had similar performance if the forward AS-level path was different.

Given the apparent importance of congruent AS-level paths in IPv4 and IPv6, we measured to what extent such congruence exists today, and how this has evolved historically. We measure IPv4 and IPv6 AS paths from seven vantage points (ACOnet/AS1853, IIJ/AS2497, NTT/AS2914, Tinet/AS3257, HE/AS6939, AT&T/AS7018, NL-BIT/AS12859) which have provided BGP data to Routeviews and RIPE RIS since 2003. The figure below plots the fraction of dual-stack paths that are identical in IPv4 and IPv6 from each vantage point over time. According to this metric, IPv6 paths are maturing slowly. In January 2004, 10-20% of paths were the same for IPv4 and IPv6; eight years later, 40-50% of paths are the same for six of the seven vantage points.

Fraction of identical dual-stack paths over time

Even though the fraction of congruent AS-level paths has been increasing, it is still only around 50%. It is interesting to study the reasons for the divergence. Is it the case that some ASes or AS links from the IPv4 graph are not present in IPv6? How strong could the congruence possibly be, given the set of AS links and ASes from the IPv4 graph that are already present in the IPv6 graph? For each link in an IPv4 AS path toward a dual-stacked origin AS, we examine whether that link is present in the IPv6 topology, regardless of the AS path on which it appears. The figure below shows that currently, 60-70% of AS paths could be link-identical in IPv4 and IPv6 without configuring a new BGP peering session, because for these paths each IPv4 link is already present in the IPv6 topology, just not yet part of an observable BGP-policy-compliant path between the edges.

Fraction of dual-stack paths that could be link-identical over time

We take a step further and examine what would happen if each IPv6-capable AS were to establish equivalent peerings in IPv6 and IPv4. For each AS in an IPv4 AS path toward a dual-stacked origin AS, we examine whether that AS is present in the IPv6 topology. The figure below shows the fraction of IPv4 AS paths where each AS on the IPv4 AS path is present in the IPv6 topology. If current IPv6-capable ASes established equivalent peerings in IPv4 and IPv6, 95% of AS paths could be node-identical in IPv4 and IPv6; that is, for an AS link on such a path, both ASes are present in the IPv6 topology, and both ASes already peer in IPv4. If these ASes also started IPv6 peering, we could see the AS paths converge.

Fraction of dual-stack paths that coluld be node-identical over time

These results are encouraging, but they are even more motivating when juxtaposed with the above performance measurements which show IPv4 and IPv6 data plane performance is comparable when the AS paths are the same. Together, these results demonstrate the undeniable benefit of BGP peering parity between IPv4 and IPv6 AS-level topologies.

Twelve Years in the Evolution of the Internet Ecosystem

Our recent study of the evolution of the Internet ecosystem over the last twelve years (1998-2010) appeared in the IEEE/ACM Transactions on Networking in October 2011. Why is the Internet an ecosystem? The Internet, commonly described as a network of networks, consists of thousands of Autonomous Systems (ASes) of different sizes, functions, and business objectives that interact to provide the end-to-end connectivity that end users experience. ASes engage in transit (or customer-provider) relations, and also in settlement-free peering relations. These relations, which appear as inter domain links in an AS topology graph, indicate the transfer of not only traffic but also economic value between ASes. The Internet AS ecosystem is highly dynamic, experiencing growth (birth of new ASes), rewiring (changes in the connectivity of existing ASes), as well as deaths (of existing ASes). The dynamics of the AS ecosystem are determined both by external business environment factors (such as the state of the global economy or the popularity of new Internet applications) and by complex incentives and objectives of each AS. Specifically, ASes attempt to optimize their utility or financial gains by dynamically changing, directly or indirectly, the ASes they interact with.

The goal of our study was to better understand this complex ecosystem, the behavior of entities that constitute it (ASes), and the nature of interactions between those entities (AS links). How has the Internet ecosystem been growing? Is growth a more significant factor than rewiring in the formation of new links? Is the population of transit providers increasing (implying diversification of the transit market) or decreasing (consolidation of the transit market)? As the Internet grows in its number of nodes and links, does the average AS-path length also increase? Which ASes engage in aggressive multihoming? Which ASes are especially active, i.e., constantly adjust their set of providers? Are there regional differences in how the Internet evolves?

To answer these questions, we analyzed AS-level topology snapshots constructed from publicly available routing tables collected at Routeviews/RIPE. We selected a series of topology snapshots spaced 3 months apart, spanning twelve years from 1998-2010. Unfortunately, the available historical datasets from RouteViews/RIPE are not sufficient to study the evolution of settlement-free peering links. So we restricted the focus of this study to the evolution of AS types and of customer-provider links. We developed a method to classify ASes (with 75-80% accuracy according to our validation efforts, detailed in our recent TON paper) into a number of types depending on their business function, using observable topological properties of those ASes. The AS types we consider are large transit providers (LTP), small transit providers (STP), content/access/hosting providers (CAHP), and enterprise networks (EC).

Our findings highlight some important trends and evidence for how these trends may play out in the future:

  • The IPv4 AS-level Internet has gone through two growth phases: an initial exponential phase up to mid/late-2001, followed by a slower exponential growth thereafter. Contrast this with the growth of the IPv6 AS-level graph, which has been growing exponentially since 2003. (More analysis comparing growth trends in the IPv4 and IPv6 topologies coming soon!)
  • The average path length, however, remains practically constant around 4 AS hops, meaning that the network densifies.
  • Currently, 81% of link births are associated with existing ASes rather than new ASes (rewiring versus growth); similarly, 86% of the link deaths are due to rewiring. This implies that most of the dynamics in the network are due to births and deaths of links between existing ASes.
  • We classified ASes according to economic considerations and business types. We find that most of the growth is due to Enterprise Customers (ECs) at the network edge. The average multihoming degree of ECs has remained roughly constant, but has increased significantly for other network types -- transit providers (both regional and global), and Content/Access/Hosting providers. The previously mentioned densification process is thus driven by transit providers and content/access/hosting providers. In terms of rewiring, CAHPs appear to be the most active, while ECs are the least active.
  • We introduced two provider metrics, attractiveness and repulsiveness, to quantify the ability of a provider to attract and retain customers. We found that both the attractiveness and repulsiveness of a provider are correlated to its customer degree. Also, many providers exhibit strong repulsiveness 3-9 months after exhibiting attractiveness, i.e., these providers attracted new customers but were unable to retain them. We define the set of providers that accounted for 70% of all customer gains across two consecutive topology snapshots as attractors, and the set of providers that accounted for 70% of all customer losses as repellers. We found that the set of attractors and repellers is increasing in size. This set of providers is dominated by those in North America and Europe; while the number of such providers in Europe is increasing, that number in other regions is relatively flat. This indicates that the market of transit providers, at least in Europe, does not seem to be consolidating.
  • With respect to regional growth, we find that the Internet market, in terms of the number of enterprise, access/hosting/content and transit networks is now larger in Europe than in North America. Additionally, since 2004-2005, a larger fraction of active customers (customers that changed their set of providers between two consecutive snapshots) are in Europe than in North America. We note that these trends refer to the number of networks in various regions, and do not reflect the size of their customer base or advertised IP address space. Our measurements hint at an increasing European influence on the Internet ecosystem.

We have made the datasets collected as part of this work publicly available. We have also developed an interactive interface to the data that allows a user to query the historical connectivity (number of customers, providers, and peers) of a set of ASes. The user can also compare the connectivity of selected ASes with the average of different types of ASes -- classfied according to their business types as Enterprise Customer, Small Transit Provider, Large Transit Provider, and Content/Access/Hosting Provider. We also allow the option of computing the degrees at the level of organizations. The user must enter an AS number belonging to the organization (e.g., 7018 for AT&T), and if a matching organization is found for that AS number, we display the connectivity of the organization as a whole. We are in the process of expanding our AS-to-organization database, so in the future the script will incorporate a larger set of AS-organization mappings.

Targeted Serendipity: the Search for Storage

On the heels of our recent press release regarding fresh publications that  make use of the UCSD Network Telescope data, we would like to take a moment to thank the institutions that have helped preserve this data over the last eight years. Though we recently received an NSF award to enable  near-real-time sharing of this data as well as improved classification, the award does not cover the cost to maintain this historic archive. At current UCSD rates, the 104.66 TiB would cost us approximately $40,000 per year to store. This does not take into account the metadata we have collected which adds roughly 20 TB to the original data.  As a result, we had spent the last several months indexing this data in preparation for deleting it forever.

Then, last month, I had the opportunity to attend the Security at the Cyberborder Workshop in Indianapolis. This workshop focused on how the NSF-funded IRNC networks might (1) capture and articulate technical and policy cybersecurity considerations related to international research network connections, and (2) capture opportunities and challenges for the those connections to foster cybersecurity research.  I did not expect to find a new benefactor for storage of our telescope data at the workshop though, in fact, I did.

During the workshop, I mentioned to the group that  we were preparing to purge historic darknet data for lack of funds to pay for storage. Upon hearing of our plans to delete  the data, a NERSC System Administrator offered to store the data in NERSC's tape archive. He understood the relevance of the data to cybersecurity research and the value of longitudinal analysis on this fairly rare and unique data type. In less than a month, we had accounts and began the work of moving the data.

CAIDA would like to thank the San Diego Supercomputer Center for archiving the UCSD Network Telescope data since 2003. The IBM HPSS  and more recently Sun SamQFS archival storage systems dutifully preserved and delivered the 100+ Terabytes of raw pcap traces we have archived over the last eight years.

We would also like to thank the National Energy Research Scientific Computing Center (NERSC) and ESnet for the resources that  allowed us to continue to preserve this data. On 22 March 2012, we started the transfer via ESnet shown in-flight in Figure 1  to  the NERSC HPSS facilities. The transfer completed in roughly one week's time and sustained an average of 1.52 Gbps limited by local host disk I/O.

120 TB Transfer from SDSC to NERSC via ESnet
Figure 1. 100+ TB transfer of UCSD Network Telescope data from SDSC to NERSC via ESnet.

Figure 2 below presents an interesting heat map visualization of the data collection volume. Each vertical bar represents one day of data, while the horizontal bands represent the size of a compressed file (in pcap format) of captured traffic for an hour of the day. We color each data point based on its deviation from the median hourly captured traffic file size. Specifically, an hour equal to the median file size we color red. Hours with (compressed) traffic volumes at twice (or more) of the median we color yellow, and hours with no data appear black. So, hotter colors mean more data. Data collection on the telescope is a best-effort service -- outages show up as vertical black bars. This plot also reveals an increase in the amount of data stored after April 2009 due to the removal of an upstream rate limit filter on incoming packets (We removed that filter in the wake of the advent of the Conficker worm, in order to study it.) The color changes in the heat map also show the diurnal variation in traffic volume, although since this type of traffic originates from most time zones of the world, the "busy hour" is not sharply delineated.

Heatmap of the UCSD Network Telscope data
Figure 2. Heatmap of the UCSD Network Telescope data.

Internet Censorship Revealed Through the Haze of Malware Pollution

We were happy to see the coverage of UCSD's press release describing two papers we recently published, introducing new methods and applications for analyzing dark net data (aka "Internet background radiation" or IBR).  The first paper, "Analysis of Country-wide Internet Outages Caused by Censorship", presented by author Alberto Dainotti last November at IMC 2011, focused on using IBR in conjunction with other data sources to reveal previously unreported aspects of the disruptions seen during the uprisings of early 2011 in Egypt and Libya. The second paper, "Extracting benefit from harm: using malware pollution to analyze the impact of political and geophysical events on the Internet", published in ACM SIGCOMM CCR (January 12), used IBR data observed by UCSD's network telescope to characterize Internet outages caused by natural disasters. In both cases the analysis of this (mostly malware-generated) background traffic contributed to our understanding of events unrelated to the malware itself. Our press release was picked up by several online publications, including The Wall Street Journal Blog, ACM TechnewsCommunications of the ACM Web siteSpacedailyPhysorgTom's GuideProduct Design & DevelopmentNewswiseDomain-bEurekAlertEurasia reviewSecurity-today.comEverything San DiegoSpacewar Cyber War.

The papers are also available on CAIDA's publications page.

Second Workshop on Internet Economics (WIE2011)

As part of our NSF-funded network research project on modeling Internet interconnection dynamics, we hosted the second Workshop on Internet Economics (WIE2011) last December 1-2. The goal of the workshop was to bring together network technology and policy researchers with providers of commercial Internet facilities and services (network operators) to further explore the common objective of framing an agenda for the emerging but empirically stunted field of Internet infrastructure economics. The final report (http://www.caida.org/publications/papers/2012/wie11_report/) attempts to capture the content, structure, and depth of the discussions, and presents relevant open research questions identified by workshop participants. From the intro (but the 5-page pdf is worth reading):

Building on the success of our first (virtual) Workshop on Internet Economics (WIE09), we expanded the scope and depth of the second workshop in this series, inviting experts in the following topics of interest: peering strategies and conflicts; content delivery; traffic and topology dynamics of the peering ecosystem; sustainable business models and industry structure. The workshop format was structured to promote constructive, focused engagement. Attendees presented research results, offered updates on data sources, moderated topic discussions, served as formal respondents to other speakers, and critiqued the relevance and potential impact of the presented results An intended goal was to establish a set of open questions that can frame an Internet economics research agenda, and more specifically to improve the realism, utility, and predictive power of economic models of Internet topology and dynamics. Even this expanded scope falls short of the recognized breadth of the emerging discipline of Internet economics; we plan to include other recurring issues and questions at the next workshop, including the economics of privacy, advertising, censorship, and intellectual property.

NASA's recent DNSSEC snafu and the checklist

Reading about NASA's recent DNSSEC snafu, and especially Comcast's impressively cogent description of what went wrong (i.e., a mishap that seems way too easy to 'hap'), I'm reminded of the page I found most interesting in The Checklist Manifesto:

We're obsessed in medicine with having great components -- the best drugs, the best devices, the best specialists -- but pay little attention to how to make them fit together well. Berwick (president of the Institute for Healthcare Improvement in Boston) notes, "Anyone who understands systems will know immediately that optimizing parts is not a good route to system excellence"... He gives the example of a famous thought experiment of trying to build the world's greatest car by assembling the world's greatest car parts. We connect the engine of a Ferrari, the brakes of a Porsche, the suspension of a BMW, the body of a Volvo. "What we get, of course, is nothing close to a great car; we get a pile of very expensive junk."

Nonetheless, in medicine that's exactly what we have done. We have a $30B/year National Institutes of Health, which has been a remarkable powerhouse of medical discoveries. But we have no National Institute of Health Systems Innovation alongside it studying how best to incorporate these discoveries into daily practice -- no NTSB equivalent swooping in to study failures the way crash investigators do, no Boeing mapping out the checklists, no agency tracking the month-to-month results.

The same can be said in numerous other fields. We don't study routine failures in teaching, in law, in government programs, in the financial industry, or elsewhere. We don't look for the patterns of our recurrent mistakes or devise and refine potential solutions for them.

But we could, and that is the ultimate point. We are all plagued by failures -- by missed subtleties, overlooked knowledge, and outright errors. For the most part, we have imagined that little can be done beyond working harder to catch the problem and clean up after them. We are not in the habit of thinking the way army pilots did as they looked upon their shiny new Model 299 bomber -- a machine so complex no one was sure human beings could fly it. They too could have decided just to "try harder" or to dismiss a crash as the failings of a "weak" pilot.

Instead they chose to accept their fallibilities. They recognized the simplicity and power of using a checklist.

And so can we. Indeed, against the complexity of the world, we must. There is no other choice. When we look closely, we recognize the same balls being dropped over and over, even by those of great ability and determination. We know the patterns. We see the costs. It's time to try something else.

Try a checklist.

The Checklist Manifesto, Atul Gawande.

More on this topic the next time it happens, which I reckon won't be too far into the future. (Cringe.)

The Menlo Report and its Companion bring ethical guidelines to ITC research

Finally, a process we started almost three years ago has reached a milestone: the first public draft of The Menlo Report: Ethical Principles Guiding Information and Communication Technology Research and its Companion Report were posted on the DHS and SRI web sites (respectively) last month.

DHS's Science and Technology Directorate, through its PREDICT program, sponsored this report on ethics in Information and Communication Technology Research (ICTR). The culmination of a multi-year effort by network and security research stakeholders to lay out a guiding framework to identify, navigate, and resolve ethical issues in ICTR, this report is intended to be a dialogue launch point for the community of researchers, oversight entities, and policymakers to reflect on ethical issues in security and network research. Public comments are encouraged via the Federal Register through 27 February 2012. I'm pretty sure all comments are responded to and/or integrated into the next version of this report. Hopefully the report will also be the topic of discussion at some conferences and workshops this year, so that the community can get out ahead of these issues before we find ourselves facing legislative overreaction to catastrophe (or even perceived catastrophe). Please consider reading and submitting a comment.

The 2nd NDN Project Retreat

I kicked off 2012 with a visit to Colorado State University in Fort Collins, CO to attend the principal investigators (PI) retreat for the Named Data Networking Project, one of four projects funded under NSF's "Future Internet Architecture" (FIA) program. Impressive progress since the first FIA meeting, with substantial development and coordination of the NDN Testbed connecting the initial participating institutions, including network status reporting, state of (phase-one) OSPF routing, and testbed status pages. This two-day meeting packed in a wide range of collaborative discussions of architecture and implementation issues, including: topology and namespace structure and constraints; organizational structure and network management; routing and forwarding strategy; security issues such as attribution and privacy; early experiences with application development; evaluation and measurement; social and ethical values in technology design; and educational outreach (classes teaching NDN concepts). We also discussed how to dispel the misconception that NDN is simply collaborative web caching. (The caching is essential but the most revolutionary piece of this new communication model is retrieving data by names.)

Those familiar with the new emerging information-centric networking movement in the computer science research community will recognize NDN's fundamental theme: replace the endpoint (identified by an IP address) as the fundamental anchor of the communications architecture with the data (identified by a name). To communicate in NDN, users post named interest(s) that propagate toward where the data resides (now relying on conventional routing protocols for the underlying routing fabric but eventually hopefully using previously developed revolutionary greedy routing mechanisms) and receives, from cryptographically vetted publishers, signed object(s) matching the requests. Conceptually simple, with many collateral benefits offered by the minimization of unnecessary layers. The application is much closer to the network. Mobility is inherent, since the notion of location has been removed as an architectural anchor.

While at least a dozen papers have resulted from this project thus far, even more tangible progress has occurred on the development and experimental deployment side. A key strength of this project (as mentioned previously) is a deployment path via a testbed overlay on the current Internet. Beichuan Zhang and Lan Wang have coordinated an OSPFN implementation that distributes name prefixes in OSPF and ccnd, and a ccnx-dhcp to help local bootstrapping, which will eventually include configuring default routes, local topology and hub discovery. Applications are already running on the NDN testbed including audio, video, and multi-user chat, which are being used by weekly project coordination calls; additional performance-related testing has been conducted using supercharged PlanetLab nodes.

In parallel, different teams are pursuing the various threads of research promised for the NSF project. Patrick Crowley is leading the investigation of how fast we can get NDN nodes to forward packets, and building traffic generators to evaluate and inform the protocol design. The security research team will present their first preliminary analysis of privacy, anonymization, and signature efficiency in NDN at this month's NDSS conference. Edmund Yeh is creating a stochastic control and optimization framework to to formally (analytically) evaluate network performance, as well as coupling theoretical and experimental evaluation of joint forwarding and caching algorithms.

One of the next big R&D challenges is effective measurement techniques, not only for network management and performance evaluation -- (``This node is being flooded with interests!'') -- but also to support new types of network routing and application development and debugging (``Why is my application not getting the data?'').

We still need to study the impact of topology structure on network operations and management as we expand the set of external participants experimenting with the current platform and applications. We also still need core management functions such as methods to identify misbehaving nodes/apps, tools for debugging, log analysis, and traffic flow, the equivalents of chargen, traceroute, mechanisms for discovering one's own local globally routable namespace (NDN prefix discovery) and other routing and institutional key information when joining a new network.

And of course, the eye of the volcano: the data namespace that NDN utlizes, including policy-relevant constraints that might determine what information should be exposed by the namespace structure. Because NDN object names may convey topological as well as content information, network elements could present treasures of performance, topology, and usage data that we can only dream about in the current architecture. But unlike today's Internet, which convolves topological and organizational (peering) structure with the Autonomous System abstraction, the NDN architecture distinguishes these functions: signatures frame organizational/peer structure, while names frame the topological structure. There are obvious and not-so-obvious implications for privacy and attribution of communications, and we devoted an entire session to discussing social values that guide design decisions, with attorney Paul Ohm promising to help us assess the strength and form of expected tussles should an NDN architecture gain deployment traction.

Colorado State (home of PIs Dan Massey and Christos Papadopoulos) did a fantastic job of hosting the meeting, including a poster session and reception the evening of the first day. Several posters described undergraduate projects in Christos' recent undergrad class on on NDN networking: running a traditional (modified) IP web traffic generator over the NDN testbed; repeating (and confirming) the 2009 CoNEXT paper experiment on PlanetLab); and a content caching study at CSU's border router (estimating how much content is static (about half by requests) vs dynamic, and redundant request patterns). The second day included lots of discussion of what applications and supporting tools we should pursue next: including a graphical name space browser; graphical PIT viewer; a serverless Twitter-like application with scope control over message distribution; and a distributed, topic-based discussion board application to facilitate collaboration.

Toward the end of the meeting we discussed NSF's request for thoughts about next steps after the FIA program currently funding this work (now half-way through its three-year budget). There are tremendous opportunities for synergy with other NSF-funded information science communities such as the Cyber-Physical Systems or the DataNets programs, to experimental deployment in production science settings such as the Open Science Grid (OSG), a national distributed computing grid for data-intensive research. Perhaps most exciting is the potential opportunity that Kevin Thompson (of NSF's Office of Cyberinfrastructure) described at Internet2's last Joint Techs meeting: in response to recently commissioned strategic advice, NSF wants to leverage successful R&D investments by transitioning them into campus environments on a broad scale, i.e., with a dedicated program. Since the NDN architecture was designed to solve many of the problems now being faced by campus networks (as well as the rest of the world), I'm optimistic that we could someday see an NDN-NSFNET. Lots of known unknowns and unknown unknowns along that path, but what an exciting path!

[Thanks to our lead PI Lixia Zhang of UCLA for help with this entry.]