Skip to main content

According to the Best Available Data

The CAIDA blog — commentary and analysis of Internet measurement, infrastructure, policy, and economics, published since 2006. Formerly hosted at blog.caida.org. See About the CAIDA Blog · Categories · All entries by month · RSS feed.

my first FCC TAC meeting

I recently attended my first FCC Technological Advisory Council meeting (video archives). A week before the meeting we received a memo from the chairman of the committee (Tom Wheeler) notifying the committee of a "clear and challenging mandate from Chairman Genachowski: to generate ideas and spur actions that lead to job creation and economic growth in the ICT [information and communication technologies] ecosystem." Specifically, "The TAC will focus on the short term implementation of innovative ideas to create investment and jobs, as opposed to long term regulatory changes."

I admit some suspicion at having the FCC's mission hijacked for short-term political ends, after they spent so much time on a fantastically visionary National Broadband Plan, but the unemployment statistics are certainly staggering, so I accept the assignment as civil duty. What seems less sensible is the implication that FCC's current mission and strategy is somehow not conducive to creating jobs. On the contrary, if the FCC kept their commitment to promote competition in the broadband Internet market (as other countries do) they would not only create jobs, but advance several other goals in the Communications Act as well as the more recent Broadband Plan.

The 39 TAC members were asked several questions to inspire discussion at the first meeting. Specifically, we were asked to suggest policies to spur ICT investment in selected "growth field" areas: spectrum, health care, energy/environment, public safety, transportation, and Devices/Internet of Things, as well as suggesting additional growth areas. The questions and my thoughts, refined based on what I heard at the meeting, are below.

  1. What (if any) growth field categories would you add to the list?
    1. Networks naturally give rise to markets. Markets generate jobs. The FCC's mission is already to promote competition and user choice, which means lowering the barriers to creating network infrastructure. We have already known for decades, if not centuries, what policies promote competition. The FCC needs to (re)-implement them. A related obvious opportunity to promote growth of network infrastructure (and thus markets, and thus jobs) is policy support for ad hoc or  mesh networking and municipal broadband projects that can take advantage of underutilized spectrum.
    2. Online education, such as Steve Midgley's Learning Registry. The FCC's Director of Education for the National Broadband Plan, Steve is now at the Department of Education executing on his vision to "provide valuable federal learning resources within a relatively simple web framework, and let others utilize our functionality to create interesting and valuable new solutions. If those solutions are effective, more educators will use them, and more innovators will develop new products for the network." Sounds like precisely the kind of jobs we need to build a labor force capable of doing the rest of the most productive jobs that emerge this century.
  2. What short term, job creating actions or policy changes immediately come to mind for discussion after reading the memo?
    1. A gap in interagency cooperation handicaps the FCC's ability to promote job-creating innovation. For example, the FCC could (and should) work with the National Science Foundation on a joint Internet measurement science program. A measurement science initiative would support the transparency objectives in the Broadband Plan, which would leverage NSF research funding as well as create the kind of jobs (graduate students!) that are not only low cost, but also generate innovation and future jobs. Instead the FCC is funding a British company to do performance testing from closed (proprietary) devices running closed (proprietary) software, an initiative that does not seem to be creating any jobs for Americans, nor doing much for transparency (though it is really too soon to tell).
    2. The FCC should also work with the Department of Homeland Security's Cybersecurity R&D center to help deploy security technologies (such as DNSSEC) in national infrastructure. The FCC and the DHS have another obvious synergy: DHS's privacy-respecting data-sharing framework which could support sharing of data required under the NTIA BTOP awards with researchers who have long needed better quantitative data to inform policy.
    3. IPv6 -- the only term uttered more than "jobs" at the TAC meeting. A broad consensus seemed to emerge that the FCC should do something, even if only use its "bully pulpit", to promote IPv6 investment and deployment. I have doubts about the necessary political will here -- the backward incompatibility between IPv4 and IPv6 means that not only will governments have to regulate IPv6 into existence, but they will ultimately have to regulate IPv4 out of existence (think digital TV transition) if IPv6 is to survive. The FCC faces quite a challenge trying to bully IPv6 into deployment; the world's need for it, despite its limitations and incentive-incompatibilities, is the best justification yet for investment in future Internet architecture research (see (4) below).
  3. In what ways can we maximize new efficiencies in technology and industry (i.e. spectrum sharing, convergence of communications with other sectors) towards job creation and new investment?

    If we make sure efficiencies are sufficiently documented, capital will naturally move toward investment in them. Increasing sunlight will do a world of good here. Examples include maintaining online maps of infrastructure, documenting how the infrastructure is provisioned and maintained, and correlating provisioning models with economic statistics including employment in the region.

  4. What new technologies are you aware of have great potential or excite you? What are the barriers to it developing to job-producing, commercial scale?
    1. Eduroam, a role model project that started in Europe, but leveraging U.S. academic networking security technology. It's likely already creating jobs, but  extending this model to other communities would both create jobs, extend broadband coverage at low cost, and improve security all in one package.
    2. Apple's iPhone/iPad and (more importantly) its more open emerging competitors. Enforcing their own principles of broadband deployment in the more pervasive mobile realm, i.e., ensuring consumers can choose which applications, services, and providers they use on such devices, the FCC will create a level playing field that will generate markets (and thus jobs) on its own.
    3. Named-Data Networking, one of four Future Internet Architecture projects funded by the National Science Foundation's Division of Computer and Network Systems. The barriers to such an ambitious vision include challenges in understanding of how technology and policy [must] co-evolve, and testbeds to study such new architectures at large scale. (Disclosure: I'm enthusiastically participating in this project, see a future blog entry.)
    4. Online education/curriculum development (see learningregistry.org above), to help develop scalable sustainable software systems to collaboratively develop academic curriculum materials.
    5. Social computational systems, another one of many exciting  programs currently coming out of the National Science Foundation.
  5. Imagine the next great technological product or service. What are the 2 or 3 technology advances that get us to that new product or service?

    This question embeds a false assumption that technology rather than policy is the current gatekeeper of ICT innovation. The increase in density, and finer granularity, of connectivity spurs current innovation. The shortest path the FCC can take to job creation requires not new, or any specific, technology, but policies that incent investment in -- and equitable access to -- ubiquitous network infrastructure, by a healthy ecosystem of competing service providers, which will enable new playing fields and inspire technologies that can leverage them. New products and services, and the jobs that provide them, will follow. Facilitating development of richer network infrastructure with symmetric bandwidth will enable all consumers to be producers -- many will create their own jobs! But current economic models and policies for infrastructure provisioning will not allow this breakthrough.

CAIDA's Annual Report for 2009

[Executive Summary from our annual report for 2009, which took longer than we expected to finish this year as we've been overly busy with material for 2010's report..]

Our current research projects span topology, routing, traffic, economics, and policy. Our infrastructure activities support several measurement-based studies of the Internet's core infrastructure, with focus on the health and integrity of the global Internet's topology, routing, addressing, and naming systems.

We made significant advances (again...) in Internet topology research, supported by the expanding Ark measurement infrastructure and growing interest in understanding more about the Internet's robustness, security, and scalability. We continue to share the largest Internet topology data sets (IPv4 and IPv6) available to academic researchers, and we share many aggregated annotated derivative data sets publicly, including rankings of ISPs annotated with (our estimated) business relationships between autonomous networks. Our topology measurement platform supports IPv6, and ten of our hosting sites provide IPv6 connectivity. We have developed substantial additional software to better support distributed measurement experiments. Specific to our IPv4 topology mapping project, we have taken on the task of optimizing and improving on existing techniques for IP address alias resolution for large Internet graphs, and are planning to package up and release an implementation of our algorithms next year. In 2009 we expanded the capability of other researchers to use the Ark infrastructure for independent experiments, including an extensive Internet-wide test of network filtering hygiene.

On the theoretical side of topology research, we finally published our topology modeling framework that treats annotations as an extended correlation profile of a network, which supports rescaling topologies while retaining the same (measured) annotation profile. We also advanced our exploration of geometric structure underlying Internet-like topologies as observed in our and other measurements. Specifically, hyperbolic geometry captures an important property of complex networks: exponential expansion in space. We explored even deeper connections between network topological structure (e.g., degree distribution, clustering) and physical phenomena such as curvature and temperature.

These discoveries about topology drive our routing research agenda, a long-term objective of which is to enable dramatically more scalable global Internet routing. We explored the ramifications of the discoveries we made last year regarding efficient routing on graph topologies statistically similar to those of the Internet. Based on the evidence, e.g, clustering, observable on the Internet and other complex networks, we found that underlying hyperbolic hidden metric spaces provide a natural explanation for why so many of these complex networks found in nature can achieve such phenomenally efficient (greedy) routing without distributing global topology knowledge. Since the distribution of global knowledge about network structure is perhaps the most critically limiting requirement of the current Internet interdomain routing system, we are still investigating theoretical details of a potentially radical solution to Internet routing scalability, which takes advantage of what nature knows that we do not (yet).

We undertook several traffic analysis activities, including creating a structured taxonomy of Internet traffic classification papers and their data sets, and analyzing the "Day in the Life of the Internet" 2009 data set, consisting of 24 hours of detailed DNS packet data collected at many participating root servers as well other high-profile DNS servers. We have reduced our traffic analysis activities in lieu of pursuing progress in the policy space through participation in DHS's PREDICT project (Protected Repository of Data for Internet Cyber Threats). As part of this project, we have proposed a more flexible privacy-sensitive data-sharing framework and an experiment to test it on the UCSD network telescope instrumentation next year.

We are growing the scope of our economics and policy research. We responded to several requests from Internet governance as well as U.S. government agencies for comments and guidance on policy matters. We launched a workshop series in Internet economics, to try to begin framing a research agenda for the emerging but stunted field of Internet infrastructure economics. On the theoretical side, we published an analytically tractable model of Internet evolution at the level of Autonomous Systems (ASes), which builds on the preferential attachment (PA) model but captures fundamental differences between transit and non-transit networks. This multi-class PA model predicts a definitive set of statistics characterizing the AS topology structure, closing the "measure-model-validate-predict" loop, and providing further evidence that preferential attachment is the main driving force behind Internet evolution.

Finally, we engaged in a variety of tool development, data-sharing, and outreach activities, including web sites, peer-reviewed papers, technical reports, presentations, blogging, animations, and workshops.

Full annual report: http://www.caida.org/home/about/annualreports/2009/

Program plan for 2010-2013: http://www.caida.org/home/about/progplan/progplan2010/

On economic frameworks for gTLDs

[I submitted the following public comment to ICANN in response to their second attempt at commissioning An Economic Framework for the Analysis of the Expansion of Generic Top-Level Domain Names. I'll link to ICANN's summary of all public comments on this report when available. -k]

This second economic report posted 16 june (pdf) is an improvement over the June 2009 reports by Dennis Carlton (pdf, pdf) but there are still too many -- and too fundamental -- flaws for it to serve as the basis of any ICANN policy on new gTLDs:

  1. Like the Carlton report, the authors still seem to think one way to "evaluate" concerns raised by others is to dismiss them without further study. George Kirikos observed one reason for the similarity between the two reports -- there was overlap in authorship. Despite the loud complaints that the previous report was not sufficiently objective, ICANN commissioned a second report that was ultimately co-authored by the same company as the first report, a fact hidden by ICANN's emphasis on only the Stanford and Berkeley co-authors in the report's description on the ICANN web site.
  2. Both reports follow the philosophy of "if we don't know what the impact might be (of expanding the number of gTLDs in the root zone), then go ahead", rather than a more conservative approach of studying how proposed changes may impact security and stability. As TBL said in his 2004 comments: "...because the DNS tree is so fundamental to the Internet applications which build on top of it, any uncertainty about the future creates immediately instability and harm."
  3. All empirical data actually analyzed in the report provides evidence that new gTLDs have not thus far achieved what ICANN claimed/expected they would do. There is no coherent logic for the report's conclusion that additional new gTLDs will have a benefit that exceeds their cost. The authors admit their empirical investigations were 'cursory examinations' at best, and seem to go out of their way to avoid quantifying effects that are amenable to more rigorous analysis, instead using vague speculative language like "might", "could", "have provided some value (to registrants)". For example, it would be helpful to have a cost analysis, albeit imprecise but with defensible (transparent) reasoning through actual numbers, for new gTLDs similar to what Bill Herrin did for routing announcements, with caveats: (See What does a BGP Route cost?).
  4. On page 39 the authors claim that:
    But if new gTLDs fail to have adequate trademark protections or if an innovative new gTLD were introduced that attracted large numbers of registrants either because it competed strongly with .com or because it reached a niche market segment that was previously underserved, then infringement rates and/or cybersquatting costs could rise significantly.

    which amounts to "if the gTLD is successful, then the costs will exceed the benefit", a rather self-defeating argument for new gTLDs in principle.

  5. The authors' argument that we should allow new gTLDs despite the overwhelming evidence that new gTLDs have never served their intended purpose because they will enable "innovative business models" begs the question -- what kind of business models, and who is going to pick winners since you also later acknowledge that ICANN will have to stop somewhere, and should only proceed incrementally, adding a few gTLDs at a time and developing methods to measure and modify policy in response to harmful impacts over time, e.g., consumer confusion.
  6. The authors propose a large number of research and data analysis projects, all of which they deem 'low priority', i.e., unimportant, mainly because such projects won't change their preliminary assumptions and preemptive conclusions, and for many of the studies they proposed, it would not be possible to get the needed data for the study anyway.
  7. The report's recommendation to "proceed incrementally" -- begs the question of how priority is decided -- and other obvious media ownership questions such as why 'ownership' of a desirable gTLD would be permanent, rather than a lease, like with spectrum? on page 61, the authors further weaken their case:
    "Because of the uncertainty surrounding the introduction of innovative new products and business models, it is difficult to analyze or predict the costs and benefits of any particular new gTLD, but one can analyze generally the expected costs and benefits of various types of new gTLDs."

    This indeed is the kind of analysis one would expect from any economic study of the impact of new gTLDs, but the authors barely scratch the surface on how such an analysis might be undertaken.

  8. At a stronger priority, the report calls for hard empirical data, with no description of which data, how such data would be shared, analyzed, used, and protected, why this data is what is needed to inform policy, and how it will do so. the report is silent on all of these more fundamental questions.
  9. George Kirikos also points out in his scathing comments (html) that the ICANN-commissioned report, despite having academic authors, seems to eschew scholarship, by failing to cite related work and how it compares to the authors' own results, and avoiding discussion of (or discounting) empirical data that sheds doubt on the wisdom of what ICANN has made clear it plans to do anyway.
  10. Similar to my observations of what's happening in the security and stability discussion of root scaling, ICANN's behavior looks like it's trying to buy rubberstamps of its current plans from commercial consultants, rather than foster what is needed in the long term: a coherent field of objective, peer-reviewed technical, policy, and economic research on Internet naming and numbering, and incentivized data-sharing to support such research.
  11. kc
    caida/ucsd

IP-AS mappings

We have performed an analysis of the IP-AS mapping obtained from Routeviews/RIPE collectors.

A crucial step in various research efforts that study the Internet infrastructure is to map an IP address to the Autonomous System (AS) to which it is assigned. The most common approach to map IP addresses to ASes is by using BGP table dumps from public repositories such as Routeviews and RIPE. We assign "ownership" of an IP address to the AS that originates the longest BGP prefix that matches the IP address. Routeviews and RIPE, however, have multiple collectors, each of which peers with a diverse set of ASes. Consequently, the IP-AS mapping obtained by using the BGP table dump from one collector could be different from that obtained from a different collector. The obvious solution is to aggregate views from as many vantage points as possible to obtain the most complete IP-AS mapping possible. In practice, however, it is common to use data from just one or two collectors, as it greatly simplifies the process of collecting and pre-processing data. The goal of our analysis is to compare different collectors, in terms of the different metrics that we are interested in, viz. address space coverage, IP-AS mapping, unique ASes, unique prefixes, unique more specific prefixes, AS links, and AS paths. Further, we study the utility of adding data from more collectors, in terms of the resulting change in the aforementioned metrics. Finally, we compare the IP-AS mapping from Routeviews and RIPE tables with that obtained from Team Cymru's whois service.

The figures above show the relative changes in address space coverage when we start with table dumps from Routeviews' LINX collector (which we choose as the base table, as it provides the largest coverage of IPv4 address space of any single table in Routeviews), and successively add data from collectors in decreasing order of address space coverage. The first plot above shows that data from additional collectors results in less than 1% increase in address space coverage, and the second plot shows that additional collectors incur a change in the IP-AS mapping for fewer than 1% of addresses represented in the tables.

We repeated the same analysis for other metrics of interest - unique prefixes, unique more-specific prefixes, unique ASes, unique origin ASes, and unique AS links. We find that additional table dumps yield fewer than 1% additional (previously unseen) ASes and origin ASes, confirming previous reports that most ASes are observable from even a few vantage points. However, for other metrics, additional table dumps matter: adding a table dump can yield up to 4.8% more AS links (shown in the above figure), up to 4.6% more prefixes and 4.7% more specific prefixes than seen in the base table. Furthermore, between 10% and 70% of the more specific prefixes seen in additional table dumps are originated by a different origin AS than in the base table.

We also compared the IP-AS mapping obtained from Routeviews/RIPE table dumps with that obtained from Team Cymru's whois service. We find that the difference between the IP-AS mapping from Cymru and that obtained by combining data from all Routeviews and RIPE collectors is small (0.7% of queried addresses returned a different origin AS). 56% of these IP-AS mismatches are due to cases where Cymru and the table dumps return a single, but different AS. A significant number (41%) of mismatches are due to Multi-Origin ASes (MOASes). In particular, 34% of IP-AS mismatches are due to MOASes where the Cymru mapping does not contain one of the ASes returned by the table dumps.

In summary, our findings are reassuring. In terms of IP-AS mapping, using data from just a few of the largest Routeviews/RIPE collectors is sufficient; adding data from more collectors does not significantly change the IP-AS mapping or coverage of IPv4 address space. Also, using data from Routeviews/RIPE is not significantly different from using Team Cymru's whois service. In fact, in the data we compared, the combination of table dumps from all Routeviews/RIPE collectors gave a better view of MOAS prefixes than Cymru's lookup service.

Growth trends in the AS-level Internet

We have studied growth trends in the number of ASes seen advertised in the global routing system from different regional registries (similar to Geoff Huston's 32-bit AS Number Report, but with per-registry trends). We used Routeviews and RIPE BGP dumps over the last 12 years, and Team Cymru's WHOIS lookup service to map ASNes to registries as of March 2010. To our knowledge, historical data to map an ASN to a regional registry at any given time in the past is not available, so we cannot account for ASN movement between registries. More information about the data collection and pre-processing is in our IMC 2008 paper, "Ten Years in the Evolution of the Internet Ecosystem" and our supplemental data page.

Our most interesting observation is that the two largest registries in terms of the number of advertised ASes (ARIN and RIPE) have shown distinctly different growth trends since 2001. Both registries showed exponential growth until mid-2001, but since then ARIN's AS count has grown linearly while RIPE's has continued to grow exponentially, though with a smaller exponent than in the pre-2001 period. The number of advertised ASes allocated from RIPE is now larger than ARIN's.

all_split

We conjecture a couple of possible reasons for the shift. One possible contribution is the presence of companies in Europe that assist enterprises in obtaining Provider Independent (PI) address space and AS numbers. RIPE Labs' Forum published a recent discussion of this issue. A second, related factor is the inclination of enterprises to seek PI address space (and ASNs) in the first place. The primary objective of enterprises in obtaining an ASN and PI address space is to multihome -- connecting to multiple upstream providers for reliability, performance, and other traffic engineering goals. A larger concentration of Internet Exchange Points (IXPs) and more competition in the European transit market could make multihoming more attractive for enterprise customers in Europe than in North America. Measurements in our IMC 2008 paper confirm that the transit market is now larger and more dynamic in Europe than in North America. But we cannot absolutely confirm this theory directly with BGP routing data, since only networks with ASNs show up in BGP, and the majority of customer ASes (both in Europe and North America) are multihomed (otherwise, technically, it should not need an ASN in the first place). It would be illuminating to examine whether enterprise customers that did not request an ASN in North America would have pursued ASN + PI address space if they were in Europe, simply because the competitive transit market in Europe makes multihoming more attractive?

data collection and reporting requirements for broadband stimulus recipients

No one was more surprised than I to see data collection requirements in the NTIA's Notice of Funds Availability (NOFA) for the Rural Utilities Service's (RUS) Broadband Initiatives Program (BIP) and the Broadband Technology Opportunities Programs (BTOP):

Awardees receiving Last Mile or Middle Mile Broadband Infrastructure grants must report, for each specific BTOP project, on the following:

  1. The terms of any interconnection agreements entered into during the reporting period;
  2. Traffic exchange relationships (e.g., peering) and terms;
  3. Broadband equipment purchases;
  4. Total and peak utilization of access links;
  5. Total and peak utilization on interconnection links to other networks;
  6. Internet protocol address utilization and IPv6 implementation;
  7. Any changes or updates to their network management practices;

Incumbents have fought hard against far less onerous data collection requirements -- indeed, the above requirements in part kept incumbents away from applying for BTOP funds. So the pragmatist in me cannot imagine these requirements actually being enforced, much less extended to existing providers of access to the public Internet as part of the national broadband plan, which the FCC owes Congress by February 17th. However, the researcher in me can imagine such requirements, in conjunction with privacy-sensitive data sharing frameworks (e.g., one we've proposed), positively transforming the state of Internet science and cybersecurity. Kudos to NTIA for this earnest attempt to improve the transparency of an industry more opaque than the financial sector (for similar reasons, and in the face of just as profound risks).

'academic' thoughts about a 'future Internet'

This post is our submitted response to NSF's call for expressions of interest in the Future Internet Architectures summit, which i am attending this week.

What scientific contributions will you bring to the discussion about Future Internet architectures?

As scientists, we are compelled to explore how the peculiar structure relates to the function(s) of complex networks. Many complex networks in nature share the peculiar structural character of the Internet, but they also manifest phenomenal behavior: they efficiently route information without any observable routing protocol overhead. This achievement is currently beyond the reach of man-made networks. The Internet still uses a 30-year old routing architecture with fundamentally unscalable overhead requirements.  Yet in those 30 years, the Internet's inter-domain topology has evolved toward a structure for which nature has superior routing technology, if only we can figure out how to use it!

The prospect of zero-overhead routing is sufficiently attractive that in our previous NeTS-FIND project we developed a new theoretical framework to study it. In our framework, nodes in real networks exist in a separate but related "hidden metric space," which guides routing without overhead or topology knowledge.  We found strong evidence that not only do hidden metric spaces underlie real complex network topologies including the Internet (http://www.caida.org/publications/papers/2008/self_similarity/), but that a greedy routing mechanism applied to such topologies and underlying spaces yields a maximum percentage of paths that successfully reach their destinations. Remarkably, these successful paths almost always are shortest, regardless of the hidden space structure (http://www.caida.org/publications/papers/2009/navigating_ultrasmall/). This explanation for why (if not how) complex networks are naturally navigable had sufficiently high interdisciplinary impact for recent publication in Nature (http://www.caida.org/publications/papers/2009/navigability_complex_networks/)

We have also developed a model of Internet growth which provides strong evidence that preferential attachment is a driving force behind Internet evolution (http://www.caida.org/publications/papers/2009/AS_evolution/).  The model yields AS-level topologies with links annotated by AS business relationships (customer-provider or peer-to-peer), and suggests that preferential attachment must be related to economic realities of ISP business decisions. To our knowledge, this is the first Internet evolution model that is realistic, parsimonious, analytically tractable, uses only measurable parameters, and "closes the loop." The last feature means that having all model parameters measured from real Internet data, and substituted in analytic solutions, we can predict peculiar structural and dynamical properties of the Internet.

Finally CAIDA contributes an active measurement platform (Ark) as well as vital data resources to the creation of an underlying discipline that formalizes our observations and understanding of large-scale, complex networked systems such as the Internet. Ark directly addresses a short-term call in the
Network Science and Engineering Council's recently published research agenda
, namely to improve the quality of measurement-driven research in the computer networking community and a broader range of scientific disciplines.  Ark provides an opportunity to test and validate hypotheses about how the current Internet operates. We are planning modules for integration with other data sources, as well as external validation of measurements and inferences against reported reality help balance the inevitable trade-off between fidelity and utility of network models.

Discuss how your research ideas might contribute to an overall network architecture, where the focus is on the system as a whole and on the interactions among the components.

We are using our current FIND funding to investigate exactly this question, including implementing and potentially deploying greedy routing over hidden metric spaces in an experimental control plane such as LISP (http://www.ietf.org/dyn/wg/charter/lisp-charter.html and http://tools.ietf.org/wg/lisp/). This implementation/deployment initiative will require a coordinated effort among different groups working on future Internet architectures.

Indeed, there are many practical technical details that still need to be worked out. Which components can we deploy incrementally? For example, we must change the semantics of IP packets to hold hidden space coordinates of the packet's destination. Since we cannot touch end systems, we need address family gateways that translate between IP and hidden space headers, similar to the mapping function implemented as part of LISP. Based on our preliminary estimates, the IPv6 header provide enough bit space to hold hidden coordinates, so that LISP does appear quite close to what we need at the control plane. However, the data plane changes are more involved.  We have begun discussions with router vendors regarding implementation constraints.

And then we still have policy, security, ownership, trust, and business models to worry about. But information dissemination (e.g., routing and forwarding) is the core function of any network. Our approach is to modernize how we understand and implement this primary function as well as the associated implications for realistic future network architectures.

What is your experience in working collaboratively in a multidisciplinary setting, across disciplines and areas of expertise as well as across academe and industry?

CAIDA has extensive experience participating in as well as coordinating and hosting interdisciplinary conversations, as illustrated by the listing of our ISMA workshop series titles (http://www.caida.org/workshops/). Many of our workshops have focused on interdisciplinary conversations explicitly structures to bridge gaps between and build connections across domains. In August 2008 we co-hosted a workshop on Networks and Navigation at the Santa Fe Institute, a unique institution dedicated to multidisciplinary collaborations on complex systems.

A strength of our recent work is the composition of researcher skills including first-hand knowledge of operational and engineering realities of the Internet, expertise in the theory and practice of Internet routing and data collection, skills in mathematical analysis and modeling of complex networks, the ability to provide realistic approximations to analytically intractable problems, experience with large-scale network simulation and emulation, and the interdisciplinary capability to broaden the impact of this project to other disciplines. The collection of results we have achieved so far demonstrate that an interdisciplinary research team can make rapid progress in formalizing our understanding of large-scale, complex networked systems.

AIMS 2009 Workshop Report

We finally posted the final report for our workshop on Active Internet Measurements (AIMS '09). The abstract:

Measuring the global Internet is a perpetually challenging task for technical as well as economic and policy reasons, which leaves scientists as well as policymakers navigating critical questions in their field with little if any empirical grounding. On February 12-13, 2009, CAIDA hosted the Workshop on Active Internet Measurements (AIMS) as part of our series of Internet Statistics and Metrics Analysis (ISMA) workshops which provide a venue for researchers, operators, and policymakers to exchange ideas and perspectives. The two-day workshop included presentations, discussion after each presentation, and breakout sessions focused on how to increase potential and mitigate limitations of active measurements in the wide area Internet. We identified relevant stakeholders who may support and/or oppose measurement, and explored how collaborative solutions might maximize the benefit of research at minimal cost. This report describes the findings of the workshop, outlines open research problems identified by participants, and concludes with recommendations that can benefit both Internet science and communications policy. Slides from workshop presentations are available at http://www.caida.org/workshops/isma/0902/.

What’s Belmont Got To Do With It?

Recently a group of Internet technology researchers, attorneys and policy professionals participated in a DHS-sponsored workshop, “Ethical Principles and Guidelines for the Protection of Human Subjects in Information and Communications Technology Network and Security Research.” Possible nickname: Belmont Flux Workshop. If you’re still glassy-eyed: (1) you have yet to engage the depths of an Institutional Review Board (IRB) in the context of network and security research; (2) you gave up after seeing “Ethical principles”; and/or (3) you think human subjects issues and network research are orthogonal.

Here's a summary of the event, and hopefully some inspiration.

The purpose of the workshop was to attempt to interpret the guidelines set forth in the three-decades-old Belmont Report as they might translate to the newer and more dynamic domain of Internet, and particularly Internet security, research. The Belmont Report was promulgated by a Commission spawned from the National Research Act of 1974. It provided guidance for protecting human subjects involved in biomedical and behavioral research supported by the now-named Dept. of Health and Human Services (HHS). This Belmont Report became the basis for HHS regulations (codified at 45 CFR part 46) which in turn became the model for the uniform rules (the “Common Rule”) for human subjects research for 14 other Federal departments and agencies.

The important takeaway from this recount of authoritative history is understanding what catalyzed it. The ground truth of our individual and collective human nature is to not take precautionary, preventative or remedial measures until we’ve been damaged, materially or otherwise. This practical truth is institutionalized in our system of law and regulation, which largely reacts to appreciable harm by proscribing and prescribing certain actions. The original Belmont Report occurred ex post facto to infamous abuses of human subjects experimentation by doctors and scientists such as in WWII concentration camps and the 1940’s Tuskegee syphilis study. As a result of these abuses, the government recognized a need to develop standards for judging doctors, scientists and researchers whose work involves human subjects. These principle-based standards have been applied in the context of formal judicial proceedings, e.g., the Nuremburg War Crime Trials, down to researchers concerned about ethically sound experiment design and review committees (e.g., IRBs) to assess whether research risks are justified.

Fast forward to today's information and communication technology (ICT) landscape, and in particular to network and security research on the global Internet, a domain that has evolved similar principles to the Belmont Report, but has no ratified method of applying them. Rather than wait for the first ‘Electronic Guantanamo Experiment’, the ultimate goal of this workshop series (there is likely to be at least another workshop) is to establish ethically defensible guidelines for current and future network and security research, so that both individually and collectively we can more effectively avoid and/or mitigate risks of harm to persons. Guidelines ratified by the research community will also help navigate the legal grey area of ICT transactions in daily operations.

To map the Belmont principles from traditional scientific disciplines into a blueprint for network and security research, we considered three axes:

  1. the boundaries between ICT network research and the accepted and routine practice of network operations management;
  2. the basic ethical principles of: (a) respect for persons (research should consider persons’ choice and opinions, should provide adequate notice and allow voluntariness, and persons with diminished autonomy deserve protection); (b) beneficence, (research should maximize possible benefits and minimize possible harms); (c) justice (benefits should accrue to those who bear any burden of the research and the burdens of the research should be distributed to the extent reasonable); and
  3. the application of those principles by way of (a) informed consent (how does it apply to different types of network measurement and experimental research?); (b) risk-benefit analysis (does the research merit the risk to subjects?); and (c) selection of subjects (are the research subjects in the same population who will benefit from the results?), respectively.

Our Game Plan:

Day 1 consisted of largely of foundational presentations to help frame the discussions of the three components above. The first panel gave background information and perspectives from Institutional Review Boards, including HHS and several academic and research organizations. The second panel was comprised of network and security researchers disclosing common and prominent scenarios that vividly illustrate the need for interpretation of these ethical principles in the expanding domain of Internet research. Finally, a few attorneys addressed prominent legal issues in empirical Internet research.

The remainder of the two-day workshop consisted of two breakout sessions, both tasked with a gap analysis between the earlier presented research data use cases and the Belmont framework, recognizing that some aspects of the framework will not translate well to the network research domain, e.g., pregnant persons being in a diminished capacity category, and other aspects will need to be added to a viable framework for network research. The case-based scenarios included: botnet research (e.g., infiltration of botnets and monitoring or disrupting traffic); wide-scale network survey research (e.g, port and wireless scanning); experiments involving reputation services (e.g., scoring and publishing blacklist data); network traffic analysis (e.g., backbone tapping, P2P research); and research involving deception of individuals (e.g., phishing research, honey-* research).

Interestingly, each group produced quite different but complementary results. One group took a high-level approach and crafted the beginnings of a fleshed out Belmont framework that could generally apply to network research, including some but not all portions of Belmont while including additional principles and application guidance.

The other group anchored off the general use cases and similarly highlighted the components of the original Belmont Report that were irrelevant and in need of interpretation. For the latter, this group expounded on specific technology, privacy and risk-assessment issues to consider.

The cost-benefit element of Belmont was arguably the most fundamental dimension of our task, and certainly the most vexing. We cannot expect otherwise, as our electronic operational lives -- both individual citizen-consumers and information age institutions -- are forming the risk-based synapses that we take for granted in traditional, analog (meatspace) activities. At the least, we are challenged to understand and frame if not (help policymakers) define boundaries of psychological, physical, legal, social and economic harms in the electronic landscape. (If you thought measuring one-way packet delay was hard..)

While few if any would argue against the benefits of empirically grounded network, security, and critical infrastructure protection research, there is little fundamental appreciation and understanding of those benefits -- and much well-founded concern regarding privacy. Ethical and legal challenges inhibit access to network data or impede in vivo network experimentation for measurement and analysis, and generalize across familiar spaces such as on-line crime, computer systems security threats, and infrastructure vulnerabilities.

Other thoughts from the workshop on how to effectively balance network research utility and ethical obligations among various stakeholders:

  1. Network researchers pursuing scientific and intellectual freedom, and empirical knowledge that will inform business models and policies predicated on economic and usage patterns, security, and social behavior, etc.;
  2. Data subjects and owners seeking the benefits of technology advancement without having to surrender control of personal information or renounce liberties and freedom of on-line movement;
  3. Network/platform owners exercising their rights in a free market economy to create wealth and cultivate business and customer relationships; and,
  4. Collective right of network and data owners to build and enhance the networks within which norms, transactions and livelihoods are maturing.

Motivated to try to get ahead of the metaphorical Milgram experiment (not to be confused with his small world experiment; we're actively trying to emulate that one in a future routing architecture) in the field of Internet research, this workshop was an initial step in that direction. I'd say we succeeded in raising the level of discourse surrounding the application of ethical principles to ICT network research and upon which specific rules may be formulated, criticized and interpreted. We'll supplement the intellectual capital we produced with subsequent workshops, dialogue and research. The eventual outcome will be a formal report (I’ll codename it “Belmont Flux Report” just for the moment) designed to serve as guiding policy for stakeholders. Stay tuned.

Where you see risks, I see opportunity. -- Alfred Blalock (performed first heart surgery, documented in Something the Lord Made).

Update: The first draft of the formal report is titled "The Menlo Report: Ethical Principles Guiding Information and Communication Technology Research"

a recent visit to the fcc

I spent a few hours at the FCC two weeks back, presented a slide version of a top ten list I wrote last year. Requested discussion topics: obstacles to data collection, how data is collected and used, policy-making based on inference, how to develop an objective knowledge base for science and policy, privacy expectations/rights versus the need for understanding the system as critical infrastructure. Audience mostly lawyers, worried about how they are going to accomplish a reasonable broadband plan. As I tried to describe in my five-minute presentation slot (and 1 slide, and more expansive blog entry) on the broadband panel at the DOC ten weeks ago, solutions begin with recognition of some underlying empirical facts, starting with one that is strangely not being emphasized by lobbyists: you can't make Wall-Street-approved margins moving bits around over long distances. Lot of implications to that reality; the sooner we admit it, the more realistic our broadband plan will be.