Skip to main content

According to the Best Available Data

The CAIDA blog — commentary and analysis of Internet measurement, infrastructure, policy, and economics, published since 2006. Formerly hosted at blog.caida.org. See About the CAIDA Blog · Categories · All entries by month · RSS feed.

what we can't measure on the Internet

As the era of the NSFnet Backbone Service came to a close in April 1995, the research community, and the U.S. public, lost the only set of publicly available statistics for a large national U.S. backbone. The transition to the commercial sector essentially eliminated the public availability of statistics and analyses that would allow scientific understanding of the Internet a macroscopic level.

In 2004 I compiled an (incomplete) list of what we generally can't measure on the Internet, from a talk I gave on our NSF-funded project correlating heterogeneous measurement data to achieve system-level analysis of Internet traffic trends:

  1. for the most part we really have no idea what's on the network
  2. can't figure out where an IP address is
  3. can't measure topology effectively in either direction, at any layer
  4. can't track the propagation of a routing update across the Internet.
  5. can't get a router to send you all available routes, just best routes (prevents realistic simulation of what-if scenarios)
  6. can't get precise one-way delay from two places on the Internet
  7. can't get an hour of packets from any backbone
  8. can't get accurate flow counts from any backbone
  9. can't get anything at all from the backbones [we used to have anonymized traces]
  10. can't get topology information from providers
  11. can't get accurate bandwidth or capacity info. not even along a path, much less per link
  12. can't trust whois registry data
  13. no general tool for `what's causing my problem now?
  14. privacy/legal issues deter research (& it was hard in a enlightened monarchy)
  15. privacy/legal issues deter measurement kc, 2004 NSF SCI PI meeting

Some caveats are in order:

  1. Although some of these phenomenon are possible to partially or imprecisely measure under certain instrumented circumstances, or within a single company, this data is not generally available for research use.
  2. There are a few small efforts underway that attempt to share existing data, e.g., PREDICT, Datapository, Datcat, Media Research Hub, but they all rely on voluntary data submissions and scant operational budgets which limits their use and impact.
  3. After 9/11, national security concerns led to an increase in measurement and access capability for law enforcement officials at both tax and consumer expense, but none of this measurement has (yet) been made available (even in anonymized form) for research use.
  4. After the telecom crash, ISPs also started to deploy more measurement capability, motivated by security concerns and perhaps even more by the need to better understand and manipulate their own traffic profiles to increase the return on their infrastructure investments.
  5. The academic network research community has (few, but loud) examples of egregiously poor judgment, e.g., deanonymizing anonymized traces without consulting those who gave you the data, violating the trust model of those who shared data, and giving providers even more reason to keep data taps closed.

So I don't mean to imply that Internet measurement is not occurring; on the contrary; it has become clear that a growing number of segments of society have access to -- and use -- sensitive private network information on individuals for purposes we might not approve of if we knew how the data was being used. But the scientific research community as well as the public remains severely under-informed regarding any macroscopic characteristics of the Internet. And although the Internet seems to survive quite well without macroscopic measurement, I also note a few reasons to worry.

  1. the growing gap between operations and scientific research, and the continuing opacity of the sector to consumers, auditors, regulators, and the public illustrates Stiglitz's information asymmetry -- the telecom bubble, crashes, restatements, and indictments of this decade are just the beginning of this systemic weakness unless the imbalance is corrected.
  2. Legislators, regulators, and politicians are engaged in deep public policy debate regarding our communications fabric, a conversations rooted in empirical questions that we cannot answer well with the current state of data availability.
  3. While the core of the Internet continues its relentless evolution, scientific measurement and modeling of its systemic characteristics has largely stalled. What little measurement is occurring reveals some disturbing realities about the ability of the Internet's architecture to serve society's needs and expectations.

It is eye-opening to note that even throughout the several decades of U.S. government stewardship of the early Internet, the only statistics collected regularly were those required by government contract. Since the privatization of the Internet in 1994-5, the United States has embraced a policy (and others have followed) that has sacrificed this data access in exchange for other public policy goals, such as Internet market expansion unfettered by the kind of regulatory reporting requirements applied to telephone companies. In fact one can attribute much of the recent industry angst to the growth success of the 90s that rendered data transport so affordable.

But Internet growth in this country has started to slow according to OECD rankings, and in particular the differentiating parameter between the U.S. and those countries ahead of us in the rankings (Denmark, Netherlands, Iceland, Korea, Switzerland, Norway, Finland, Sweden, Canada, Belgium, UK, Luxembourg, France, and Japan) has been government policy, specifically regulations governing cooperative shared use of critical communication facilties.

So now, in addition to the data/science crisis inside the ivory tower, we have set of public policy crises out in the real world: how to most cost-effectively improve -- and measure -- high-speed access to the Internet for Americans? Incumbent duopolists promise that their proprietary QoS innovations will help, but they want to charge a heavy price: not sharing infrastructure facilities. That is, the proposed solution of the incumbent telco and cablecos is to take the United States in the opposite policy direction from every nation with greater broadband penetration than we have, in order to achieve greater broadband in the U.S. And they want us to accept this strategy with no empirical data from their networks upon which to base a discussion. This level of discourse makes the prospect of regulation seem less surprising, even less disconcerting, to those seeking a healthy competitive network environment.

Of course, the first question that comes up in the discussion of broadband penetration and growth is: what and how do we measure this? And it turns out that no one is happy with how the U.S. FCC measures broadband -- not even the FCC. My goodness, what a long road we have ahead of us.

k.

It is fair to say that we need a new routing system

i get this question a lot:

at the current churn rate/ratio, at what size does the FIB need to be before it will not converge? (also sometimes pronounced 'when will the current Internet routing architecture break?')

a good question, has been asked many times, and afaik no one has provided any empirically grounded answer.

a few realities hinder our ability to answer this question.

  1. there are technology factors we can't predict, e.g., moore's law effects on hardware development.
  2. there are economics and policy and social factors we can't predict, e.g., how much convergence-capable hardware will providers/vendors be able to afford, how those costs will affect consumer prices, how that will affect consumer uptake, network growth, and industry dynamics, how regulation affects all of the above.
  3. we have no data from providers on the dynamics of BGP and IGP interactions, much less network wide convergence, so the research community can't provide any empirically grounded input into an answer.

note, however, that like the 'when do we run out of address space?' question, uncertainties in both technology progress and human behavior render any prediction of an actual convergence apocalypse timestamp rather sketchy, and i reckon someone with an agenda could devise parameters and 'observe correlations' to match their agenda.

also note that this does not mean we don't have a problem, just like not having a validated ipv4 address exhaustion timestamp means does not mean we don't have a problem with address exhaustion.

the reason we know we have a problem, and that it's only a matter of time before we'll need another approach to routing, is that the current system is inherently not scalable indefinitely, and in particular is inherently a poor fit to the topology and traffic engineering practices that underlie the 'natural' operations and evolution of the infrastructure.

this is why the IAB still has workshops about the issue even though they don't actually have any empirical data in the workshop report, and whenever the report touches on this question 'how long do we have?', they add "Editor's note: This is an area of much controversy/debate, so further investigation/community input is required" (those type of words are in the report many times, sometimes before and after the same paragraph (see section 4 on the scaling problem.)):

http://tools.ietf.org/group/iab/draft-iab-raws-report/draft-iab-raws-report-02.txt

with neither a 'macroscopic data analysis' directorate of the IETF (or IRTF) nor an industry structure that could give rise to such an activity, the IAB punts on the 'supporting empirical data' aspect of the issue, and instead focuses on what it can contribute: engineers discussing/ establishing/documenting what we do know about 'fundamental problems w scalability and proposed engineering approaches to solving them'. if you were at the nov06 ietf plenary when the IAB presented this workshop summary, you may recall a few people in the audience got up and said 'what data are you even basing this sky-is-falling stuff on?' and the IAB again acknowledged the data gap, said "we'll get back to you" and afaik at no point did they provide any data. (if they did please let me know, we'll publish the numbers..)

we're not alone in wanting better quantitative data on this topic, but such data is not essential to recognizing or understanding the problem. better data would assist those trying to get attention and resources invested in a better routing system, but that's (we are) a small and highly unprofitable market segment for those who have the data. i'm not giving up on the data challenge, but, in the meantime, it is fair to say that we need a new routing system.

k

The Future of the Internet: Q&A with kc claffy

A reprint of a recent interview of kc claffy posted by the San Diego Supercomputer Center regarding the future of the Internet:

kc claffy has played a leading role in Internet research for more than a decade. She is the principal investigator for the Cooperative Association for Internet Data Analysis (CAIDA) which is based at SDSC and provides tools and analyses to promote a robust, scalable global Internet infrastructure. As a research scientist at SDSC her research interests include the collection, analysis, and visualization of workload, routing, topology, performance, and economic data on the Internet. She has been at SDSC since 1991 and holds a Ph.D. in Computer Science from UC San Diego.


Q: You co-founded the Cooperative Association for Internet Data Analysis, CAIDA, a little over 10 years ago. Can you tell us how CAIDA has evolved, and what you’re focusing on today?

I founded CAIDA to address a problem that began the year I earned my Ph.D., and which has now grown into a crisis, despite our and others’ best efforts -- the lack of available empirical data on the public Internet as the infrastructure has privatized.

I finished my Ph.D. in computer science and engineering at UCSD in 1994, with a thesis that relied on network traffic and performance measurements from the NSFNET Backbone network, the general purpose backbone supporting the U.S. research and education community at the time. (SDSC was a transit node on NSFNET, and Hans-Werner Braun, an architect of the NSFNET, had just arrived at SDSC, so we had access to a lot of data and knowledge about the infrastructure). Under the terms of their cooperative agreement to provide the backbone, the operator Merit made certain measurements available by ftp every month, and we also took packet header trace measurements at SDSC, UCSD, and NCSA, which gave my thesis, “Internet traffic characterization,” strong empirical grounding.

The year after my graduation, the NSFNET backbone was decommissioned as commercial provisioning of Internet service by private sector players took off. Publicly available traffic data on or about large-scale IP networks also went away that year, which made me fear for the future of Internet research. I started CAIDA to try to narrow the already growing gap between the Internet research community and the Internet providers and users.

Together we managed to grow and maintain a research group through increasingly difficult science funding periods, but over the last decade our ability to get data from commercial Internet providers gradually diminished, as did the quality of science in the field.

Q: Can you explain why the ability to get data diminished?

The data that exists within commercial ISPs (Internet Service Providers) is considered proprietary. Providers worry that competitors could use it to steal customers or otherwise harm their business. Other important data is not collected at all, because there is no economic incentive to do so or any regulations requiring it. Metrics that are currently grounded in dangerously insubstantial measurement include the amounts and patterns of data traffic, the structure and evolution of Internet topology, the extent and locations of congestion, the amount or number of sources of spam, phishing, or DOS (Denial of Service) attacks, patterns and distribution of ISP interconnectivity, and other metrics that are critical to analyzing the security, stability, scalability, and sustainability of the Internet.

With mixed success, CAIDA and many others in the research community have navigated the many obstacles to collection and analysis of traffic data on the commercial Internet -- not only the technical and engineering challenges but also the more daunting legal (privacy), logistical, and proprietary considerations. But the unfortunate reality is that while the Internet has already become critical communications infrastructure for business, education, public safety, health care, and civil society, there is amazingly little rigorous empirical inquiry to inform opinion, much less policy, on how to solve problems of the Internet that have persistently resisted solution for the last decade.

Q: What have we learned from and about Internet measurement, or the lack thereof?

Over time it became clear to me that there are a common set of operational problems across the Internet industry which can be classified into four dimensions of the Internet as emerging critical infrastructure, these are safety, scalability, sustainability, and stewardship. The bad news is that making progress on all of these operational problems, even those that seem technical in nature, is blocked on non-technical issues of economics, ownership, and trust. For ten years CAIDA sought to tackle one problem -- measurement -- whose biggest obstacles had long clearly been economic (cost of instrumentation and data management), ownership (legal access to data), and trust (privacy and security obstacles to measurement). A more recent, and more painful, insight was that measurement is not unique in this regard, and that all persistently unsolved operational problems of the Internet are similarly blocked on issues of economics, ownership, and trust.

The economic forces of the industry are a key factor because they drive the policy conversations in Washington right now. Without directly confronting the economic constraints that network infrastructure providers face, the integrity of network science, communications policy, or indeed, our own national information infrastructure will always be suspect. Although emerging as the essential communications fabric of our professional and personal lives, the Internet has not yet stabilized from the tremendous privatization and commercialization of infrastructure that began in the early to mid-1990s. After a decade of boom and bust, consolidation continues, with the largest of the remaining providers publicly insisting that they will not be able to make the required investment to build out broadband infrastructure unless they can have more flexible pricing strategies to recover costs, that is, they want to implement differential pricing by type of traffic. Legal scholars have long argued that this development is a constitutional threat to the First Amendment, since providers would thus have a lever to control how users of their infrastructure communicate.

While such dramatic developments are occurring inside the policy realm of Washington and around the world, the network science community has to sit by, frustrated at being unable to engage in empirical network investigations that would support not only the scientific and engineering community but also the policymaking community, where lack of data now carries with it ominous Constitutional implications. The recent controversy over the NSA’s access to commercial Internet links only heightened the already well-established paranoia about traffic data collection, further hampering this already stunted field. At this rate, by the end of the decade the network research community will be one of the few groups of people who do not have access to Internet data!

This recognition has led to a change in strategy for CAIDA. It is no longer appropriate to pursue solutions to the Internet’s problems without tackling the related economic, ownership, and trust issues. CAIDA’s activities have always spanned the four Ss -- security, scalability, sustainability, and stewardship -- but we have begun to refocus current projects and pursue new ones that openly navigate links between technology, economics, and policy.

Q: You’ve pointed out that the lack of available measurement data on the public Internet as the infrastructure has privatized makes it hard to understand the complexities of the Internet or develop informed policy. Can you tell us why this is important?

Well, we should recognize the reality: the United States is facing a worsening information infrastructure crisis -- over the past half-decade the U.S. has fallen behind a growing list of industrialized nations in delivery speeds, price per megabit, broadband penetration rates, and other facets of broadband service provision. Our personal and national security realities are even more disturbing, since the best (but not good) available data shows a formidable profusion in the number and extent of unwanted and malicious traffic, things like DOS attacks, identity theft, spam, phishing, viruses, and worms. The more of our lives we migrate over to this digital realm, the more risk we assume. A targeted attack, relying on technology as well as social capabilities that have already been demonstrated, could cut off, for some period of time, not only our channels of personal communication and entertainment but also our banking, financial services, e-commerce, and supply chain infrastructure, creating devastating economic impacts.

Emphasizing my earlier point, regulatory, political, and market constraints on providers have rendered Internet researchers incapable of studying mission-critical aspects of the Internet and the state of its current robustness, capacity, usage, and vulnerabilities. Potential solutions to persistently unsolved problems thus remain an area of uninformed conjecture rather than rigorous, empirically grounded analysis.

Q: There’s a lot of concern about the future of the Internet and whether it will be turned into a private toll road or remain an open public information highway. What do you see as the principal opportunities, and the main challenges, that lie ahead, and how can CAIDA help?

The good news is that there’s a growing realization in society that the Internet is critical infrastructure for our nation and the world. Historically, new transport infrastructure such as railroads, telegraphs, the electric grid, started out like the Internet did -- “in the wild” and largely unregulated. Once everyone -- especially voters -- considers Internet access critical to their lives, their elected representatives will take an interest in ensuring stability and universal access as essential services. In fact, our broader reliance on the Internet has already led to discussion in the U.S. Congress and elsewhere about how the Internet should develop. For example, what requirements and incentives should there be to ensure connectivity for the significant still-unconnected segment of our own country’s population? The discussion is healthy, but the dearth of empirical data hinders informed debate.

And there are lots of policy issues at stake now. In contrast to other countries, the U.S. recently removed the policy that required open access for competitors to the pipes into people’s homes and businesses. There is no clear path to competition without open access requirements; facilities-based competition (assuming sufficient competition will emerge across entirely independent physical facilities such as DSL, cable, and satellite) has failed to fulfill its promise of recapturing U.S. leadership in the Internet industry -- on the contrary, since removal of open access the best available data suggests a drop in competition as well as -- arguably related -- in our international ranking in broadband penetration. We hope, and try to help, governments base public policy strategies on the best available empirical data, and to quantitatively measure the performance of those strategies against intended results. Having good data is essential to good policy.

Q. What do you enjoy doing outside of work?

I enjoy spending time with my family, who mostly live on the East coast unfortunately, and I enjoy music and cycling, reading, writing, playing on the Internet, and anything with my sweetheart.

The (un)Economic Internet

IEEE published this announcement of a new series of papers related to Internet economics in its may issue:
http://www.caida.org/publications/papers/2007/ieeecon/
MAY - JUNE 2007 1089-7801/07/$25.00 c 2007 IEEE Published by the IEEE Computer Society 53 Internet Economics Track Editors: Scott Bradner - sob@harvard.edu kc claffy - kc@caida.org kc claffy and Sascha D. Meinrath Cooperative Association for Internet Data Analysis Scott O. Bradner Harvard University

The (un)Economic Internet?

The Internet Economics track will address how economic and policy issues relate to the emergence of the Internet as critical infrastructure. Here, the authors provide a historical overview of internetworking, identifying key transitions that have contributed to the Internet's development and penetration. Its core architecture wasn't designed to serve as critical communications infrastructure for society; rather, the infrastructure developed far beyond the expectations of the original funding agencies, architects, developers, and early users. The incongruence between the Internet's underlying architecture and society's current use and expectations of it means we can no longer study Internet technology in isolation from the political and economic context in which it is deployed.

This article kicks off IC's new series on policy, regulatory, and business-model issues relating to the Internet and its economic viability. These articles will explore a range of topics shaping both today's Internet and the discourse in legislatures and deliberative bodies at the local, state, national, and international levels in pursuit of enlightened stewardship of the Internet in the future.

Mindful of Internet connectivity's fundamental import for advanced as well as emerging economies and its day-to- day irrelevance for the unconnected vast majority of human beings, pieces for this series will cover technology as well as political, economic, social, and historical issues relevant to IC's international readership. In this inaugural article, we provide a historical overview of internetworking and identify topics that need further exploration - topics we particularly encourage authors to cover in future articles in this series.

A History of Internet (un)Economics

The modern Internet began as a relatively restricted US government-funded research network. One of the most revolutionary incarnations of this network, the early ARPANET, was limited in scope - at its peak, it provided data connectivity for roughly 100 universities and government research sites. In the decades since, a few key transitions have been critical in radically transforming this communications medium. One of the most important of these critical junctures occurred in 1983, when the ARPANET switched from the Network Control Program (NCP) to the (now ubiquitous) Transmission Control Protocol and Internet Protocol (TCP/IP). This switch helped change the ARPANET's basic architectural concept from a single specialized infrastructure built and operated by a single organization to the "network of networks" we know today. Dave Clark discusses this architectural shift in his 1988 Computer Communications Review paper, "The Design Philosophy of the DARPA Internet Protocols." He wrote that the top-level goal for TCP/IP was "to develop an effective technique for multiplexed utilization of existing interconnected networks."

During this same period, network developers chose to support data connectivity across multiple diverse networks using gateways (now called routers) as the network-interconnection points. Preceding communications networks, such as the telephone system, used circuit switching, allocating an exclusive path or circuit with a predefined capacity across the network for the duration of its use, regardless of whether it efficiently used the circuit capacity. Breaking with traditional circuitswitching network design, early internetworking focused on packet switching as the core transport mechanism, facilitating far more economically as well as technically efficient multiplexing of existing networking resources. In packet-switching networks, nonexclusive access to circuits is normative (although companies still sometimes buy dedicated lines to run the packet traffic over); thus, no specific capacity is granted for specific applications or users. Instead, data is commingled with packet delivery occurring on a "best effort" basis. Each carrier is expected to do its best to ensure that packets get delivered to their designated recipients, but no guarantee exists that a particular user will be able to achieve any particular end-to-end capacity. In packet-switching networks, capacity is more probability-based than statically guaranteed. Internet data transport's best-effort nature has caused growing tension in regulatory and traditional telephony circles. Likewise, as the Internet becomes an increasingly critical communications infrastructure for business, education, democratic discourse, and civil society in general, the need to systematically analyze core functionality and potential problem areas becomes progressively more important.

Early developers couldn't have foreseen the level to which the Internet and private networks using Internet technologies have displaced other telecommunications infrastructures. It wasn't until the mid 1990s that visionaries such as Hans- Werner Braun started warning protocol developers that they needed to view the future Internet as a global telecommunications system that would support essentially all computer-mediated communications. This view was eerily prescient, yet core Internet protocols haven't evolved to meet increasing demands and are essentially the same as they were in the late 1980s.

A growing number of researchers are convinced that without significant improvements and upgrades, the Internet might be facing serious challenges that could undermine its future viability. Features such as network-based security, detailed accounting, and reliable quality-of-service (QoS) control mechanisms are all under exploration to help alleviate perceived problems. In response to these concerns, the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) Next Generation Networks study group (NGN; www.itu.int/ITU-T/ngn/) is working to define a very different set of protocols that would include these and other features.

Security: Not the Network's Job

Various people have offered explanations regarding the lack of security protocols in the Internet's initial design. Clark's seminal paper doesn't mention security, nor does the protocol specification for IP. Because the network itself doesn't contain security support, the onus has fallen to those who manage individual computers connected to the Internet, to network operators to protect Internet-connected hosts and servers, and to ISP operators to protect their routers and other infrastructure services. Services such as user or end-system authentication, data-integrity verification, and encryption weren't built into the core Internet protocols, so they're now layered on an infrastructure that isn't intrinsically secure. Currently, few existing studies examine the potential economic rationale for this current and continuing state of affairs and the ramifications for the infrastructure's efficiency, performance, and sustainability.

QoS:Too Easy to Go Without

The original IP packet header included a type of service field to be used as "an indication of the abstract parameters of the quality of service desired." This field, later updated by Differentiated Services, can define priority or special handling of some traffic in some enterprise networks and within some ISP networks, but it's never seen significant deployment as a way to provide QoS across the public Internet. Thus, the QoS a user gets from the Internet is typically the result of ISP design and provisioning decisions rather than any differential handling of different traffic types. Thus far, "throwing bandwidth at the problem" has proven to be a far more cost-effective method for achieving good quality than introducing QoS controls.

Yet, what happens if conditions change so that overprovisioning is no longer a panacea? The dayto- day quality most users experience from their broadband Internet service is good enough, for example, to enable voice-over-IP (VoIP) services such as Skype and Vonage, which compete favorably with plain old telephone services. However, the projected explosive growth of video and other high-bandwidth applications might increase congestion on parts of the current infrastructure to the point that special QoS mechanisms could be required to maintain usable performance of even the most basic services.

To read the rest of the paper, "The (un)Economic Internet" view:
http://www.caida.org/publications/papers/2007/ieeecon/ieeecon.xml

CAIDA's Annual Report for 2006

CAIDA's 2006 Annual Report covers last year's efforts, summarizing highlights from our research, infrastructure, and outreach activities. Our current research projects, primarily funded by the U.S. National Science Foundation (NSF), include several measurement-based studies of the Internet's core infrastructure, with focus on the health and integrity of the global Internet's topology, routing, addressing, and naming system. Our infrastructure activities, funded by NSF and DHS as well as other government and industry sources, include building a catalog of Internet measurement data sets, contributing to the (DHS-funded) PREDICT repository of datasets to support the (U.S.-based) network research community, and developing and deploying active and passive measurement infrastructure that cost-effectively supports the global Internet research community. We also lead and participate in tool development to support measurement, analysis, indexing, and dissemination of data from operational global Internet infrastructure. Finally, we engage in a variety of outreach activities, including web sites, peer-reviewed papers, technical reports, presentations, blogging, animations, and workshops. CAIDA's program plan for 2007-2010 will be available in July 2007.

Background and recent history

For just over ten years CAIDA has undertaken various approaches to narrowing a gap that now impedes the field of network research as well as telecommunications policy: a dearth of available empirical data on the public Internet since the infrastructure privatized. We have managed to grow and maintain a research group through increasingly difficult science funding periods, yet over the last decade our ability to get data from the commercial Internet has gradually diminished, as has the quality of science in the field of Internet research.

As an Internet data analysis and research group largely supported with public funding to apply measurement and analysis toward understanding and solving globally relevant Internet engineering problems, we accept a responsibility to seek, analyze, and communicate the salient features of the best available data about the Internet. And of course, we want better answers to the oft-asked question by graduate students: ``What Internet research problems are important to work on?'?

In 2003, frustrated with both the lack of access to data and the lack of progress in the Internet research community on a number of important problems in the last ten years, we began a survey to find answers to three questions: (1) What are the top operational and engineering problems in the Internet?; (2) Are we making measurable progress toward solving them?; (3) For problems that we're not making progress on, what is blocking progress?

These problems span four dimensions of the Internet as emerging critical infrastructure: safety, scability, sustainability, and stewardship. Measurement made the list of problems, no surprise there. At CAIDA we have experienced how difficult measurement of operational infrastructure has been, but the biggest problems with measurement had long been economic (cost of instrumentation and data management), ownership (legal access to data), and trust (privacy and security obstacles to measurement). More surprising was that measurement was not unique in this regard: all persistently unsolved operational problems of the Internet are similarly blocked on issues of economics, ownership, and trust (EOT).

Expansion of research agenda into policy and economics

This lesson meant a change in strategy for CAIDA. We recognize it is no longer appropriate to pursue solutions to the Internet's problems without tackling the EOT issues. CAIDA's activities have spanned the four S's (security, scalability, sustainability, stewardship) for years, but we have begun to re-focus current projects, and pursue new projects, that specifically recognize and openly navigate the economic and policy issues via activities that cross the boundaries between technology and policy. The table below lists examples of how we have expanded our horizons in the last couple of years. Most current projects will continue into 2007.

In addition to we expanded our efforts to
collecting topology for use in modeling and simulation activities, use the same data to provide a current ranking Internet service providers according to sample observations of AS topology coverage ( ASrank).
launching DatCat, for use by researchers to index data sets into a common catalog, reach out to policymakers to show them how such a catalog could support informed policy making.
upgrading our measurement infrastructure and data curation tools to keep pace with technology changes and research community needs, propose a project to provide otherwise non-existent incentive for operational networks to use these tools to contribute network meas urements to the research community.
simulating routing protocols that could scale to billions of nodes, simulate IPv4 address consumption scenarios for ARIN to support immediate policy needs.
providing the largest simultaneous collection of traffic to DNS root servers in support of the research community, do macroscopic surveys of the health of the current DNS, and participate on two of ICANN's subcommittees (SSAC and RSSAC).

If you have comments or questions, please do not hesitate to contact us at info at caida dot org.

Follow this link to CAIDA's Annual Report for 2006.

Following Up On 'A Day in the Life of the Internet' Challenge

[okay, that took about four times as long as i'd hoped, but we're done with a preliminary cataloging of the data collected for our "Day in the Life of the Internet" experiment for 2007. -k]

As a refresher, this is a follow-up to our last year's announcement that we would try out this experiment recommended by a National Academy of Sciences workshop, specifically, to capture 'a day in the life of the Internet' (DITL) to support the needs of network research. We believe the research community now has more measurement data (indexed!) than ever before about a single day of the Internet, and while the data situation is still pretty bleak, a little data is better than even less. In terms of measurements executed, we did significantly better than our practice DITL run in 2006, so there is cause for optimism about the future of this kind of experiment. As the summary makes clear, this year's progress was mostly due to contributions from outside the U.S., in particular from Korea and Japan, countries which have generally more successfully navigated data sharing issues for their research communities than the U.S. has. We are sorry to say we did not index a single trace from a commercial provider link this year, although we were pleased to get participation from 5 of the 13 root nameserver anycast cluster operators, up from 3 last year.

Access to the data collected during the DITL experiments varies based upon the policies of each collecting organization and the intended use of the data. We have indexed the datasets from participating DITL data providers in CAIDA's DatCat. We recognize that data ownership policy issues are one of the biggest problems in the Internet research community, and we encourage feedback by email on any problems researchers have obtaining specific data sets, with as much detail as possible to help us improve the utility of this experiment in the future.

We should note that in the ten years that CAIDA has tried to tackle the measurement problem, the biggest obstacles have consistently been:

  1. economic (cost of instrumentation and data management)
  2. ownership (legal access to data), and
  3. trust (privacy and security obstacles to measurement).

It has become clear that if we do not find ways to tackle the economic and policy problems of Internet measurement, the incentive for technology investment in Internet research will weaken, as will our understanding of the Internet.

Detailed Summary of the January 9-10, 2007 Collection Event

If You Can't Measure It, You Can't Manage It

The following is an excerpt from a discussion forum for the Future of the Internet Workshop hosted by the Organisation for Economic Co-operation and Development (OECD), entitled "If you Can't Measure It, You Can't Manage It". Tom Vest and KC Claffy, Cooperative Association for Internet Data Analysis (CAIDA).

1. The Internet is now a critical infrastructure and a global platform for communication and commerce. What should be the role of governments in its development and management?

The Internet's ascension to the status of critical infrastructure in no way diminishes its role as a key driver of economic productivity and development, as well as human creativity and empowerment. Accordingly, while the role of government in Internet-related matters is certain to grow and evolve, these changes should reflect the greatest possible sensitivity to both sources of the Internet's relevance.

To date the Internet has largely evolved "in the wild," in market spaces created by earlier regulatory actions (or in some cases regulatory forebearances) that were designed to shelter the embryonic Internet sector from legacy rules and dominant industry players that defined the pre-Internet era. Those pro-Internet policies gave rise to tremendous growth, dynamism, and innovation, but also rendered the entire sector opaque, unamenable to objective empirical macroscopic analysis. An Internet that qualifies as critical infrastructure should possess a measure of resilience and "survivability" sufficient to inspire confidence, at both the system-wide and regional/component levels. Such confidence is only possible when information regarding the state of Internet operations is accessible to at least a subset of neutral observers.

"If you Can't Measure It, You Can't Manage It"

Before government can consider taking any more substantive role in Internet development and management, it should take a lesson from the maxim above, which is common among commercial Internet operators. Government actors must undertake a crash course in learning how Internet technology works. They should take steps to assure that consistent and accurate data about the Internet's essential features are collected on an ongoing basis, and they should support the development of tools and talent to make sense of new "health indicators" of this rapidly growing critical information infrastructure.

2. The Internet is challenging existing business models. How can we ensure there is sufficient investment to meet the network capacity demands of new applications and of an expanding base of users?

The "economics of light" that governs capacity building in the Internet's core defies conventional assumptions about the relationship between investment, supply, and demand. Given an unbroken terrestrial right-of-way between any two points, an initial if sizable one-time capital outlay will yield what amounts to infinite network capacity between those points relative to all conceivable human demand for the imaginable future. Beyond that initial investment, additional orders of magnitude of network capacity can be added at marginal additional cost, effectively precluding the long-term economic viability of a second facilities platform on that route, perhaps forever. The Internet and telecoms crash of 2000-2001 was not caused by "irrational exuberance" or over- investment so much as by the emergence of this new economics, which is the product of fabulous and largely "undiscounted" advances in optical multiplexing technologies over the past decade. These economies, operating in the presence of open competition, have pushed wholesale networking costs close to the level of cost of service delivery.

Substantial new investment will be required to extend the scope of these economies to individual consumer premises -- the only area of significance where the issue of investment incentives *might* rise to the level of public interest. Once the one-time investment is made and the optical infrastructures are in place, there will be no material/technical difference between wholesale and retail Internet sectors. Such a development could have profound positive impact on productivity, innovation, and economic growth -- if the "economics of light" are permitted to follow the light path itself, down to the consumer level. Herein lies the fundamental paradox. If the economics are allowed to filter down, the direct returns to the facilities investment will be modest -- perhaps in keeping with other utilities investments -- but the public interest in the investment will be profound. Conversely, if the economics are allowed to be fully internalized by the facilities provider, public interests will be positively, and indefinitely, harmed.

The history of telecommunications regulation in the US and elsewhere suggests that public willingness to abide any rationing or denial of new products and services by a dominant market actor is likely to be short-lived. The simultaneous existence of other national markets where the so-called "incentive problem" has been solved will provide a constant irritant to consumers who are being actively deprived -- and a steady reminder of the price that such market failures carry in terms of job creation, economic productivity, and technological advancement at the national level.

Market power is fungible, and more often than not is used to avoid both investment and subsequent regulation. OECD regulators and commercial service providers alike should recognize the fundamental long-term instability of Internet service markets that attempt to build on (and thereby encourage) significant market power as a mechanism for promoting investment in new network facilities. Instead, all interested parties should work toward a solution that encourages the required near-term facilities investments while preserving long-term market stability. A wide variety of solutions are possible, ranging from investment incentives matched with wholesale/resale service requirements, all the way to some version of the "build-own-transfer" model of capital investment that is common in some developing economies.

No possible solution proposed here, in the absence of collective deliberation by the key stakeholders, is likely to sound persuasive -- all the more reason for discussion to continue, in full light of the long-term as well as short-term interests in Internet and economic development.

3. Innovation is taking place at the edges of the network. How do we ensure that this continues and how can it be enhanced?

Innovation is not taking place at the edges of the voice telephony network, nor at the ends of the cable or broadcast television networks. Innovation cannot happen in those contexts -- not due to any particular deficits in technology, but rather because for contingent reasons those network platforms never internalized the kind of commercial and institutional arrangements that permit experimentation and novel reuse of the service platform by anyone other than the network owner.

The Internet began as a decentralized collaborative venture, lightly mediated by a single, benevolently indifferent upstream provider. In the years since privatization, the pace of Internet growth, development, and innovation has varied phenomenally. In regions where market and regulatory factors have recreated that decentralized, collaborative environment, the Internet has thrived. Wherever Internet service delivery exists only as a side venture of a centralized national service provider, Internet development has been markedly slower and less innovative.

Given the proper conditions, innovation at the edge is probably a spontaneous phenomenon -- enhancing it is probably unnecessary, trying to stop it would probably be more difficult. However, absent the proper market conditions, innovation is no more likely to happen than it is at the edges of the cable television network, regardless of the side incentives. If maintaining and strengthening edge innovation is important, then attention to such market fundamentals is likely to be the best, and perhaps the only, place to start.

4. The Internet is perceived as not being secure, nor does it protect privacy. What steps should be taken to improve security and privacy and by whom?

Perceptions about the Internet often swing wildly on many issues, in part due to the absence of consistent benchmark data on many aspects of Internet behavior. Somewhat ironically, this absence of data attests to the successful preservation of security and privacy, for Internet operational data, by commercial network operators. Some standardized industry-wide reporting of security- relevant information might help to reduce unfounded concerns, and more importantly to clearly illuminate those areas where security and privacy vulnerabilities are most egregious. Again, if you can't measure it, you can't manage it, or set enlightened public policy for it.

More directly, additional support could be given to ongoing efforts to secure the distributed transactions that core Internet technologies employ to transmit critical source and destination information. Already implemented at the DNS level in some markets (e.g., Sweden), parallel efforts to secure IP address origination should be accelerated, and deployment of both technologies should be actively encouraged by all stakeholders.

Once this distributed, transaction-level security has been achieved, the weakest point in the Internet security chain will be at the top, where network identities are matched with real-world identities by the institutional stewards of the Internet's unique naming and protocol number resources. Recognizing this, the registries should redouble their efforts to standardize, regularize, and secure their "whois" records, and to fill in any remaining historical gaps in these critical identity records.

5. Ubiquitous networks are being deployed. What are the drivers of these developments? What will be the impacts on individuals and society?

Current demand drivers for the proliferation of ubiquitous network- enabled devices include national security and property security interests and applications such as inventory control and personnel tracking. In the near future, such devices may also empower a variety of unique and useful new consumer services. However, even in the most benign configurations, these should probably be regarded as dual use technologies whose security implications should not be forgotten or ignored. Given the potential commercial value of pervasive post- purchase consumer behavioral monitoring, societies may have to develop both the technical means to define ubiquitous network "no go zones" (e.g., disable auto-locating services on a cell phone), and perhaps also to expand the definition of personal intellectual property to include information about one's own usage of purchased goods. National authorities should anticipate these conflicts, and begin to contemplate possible responses to the proliferation of opaque "click wrap" style agreements -- now common in software -- to other consumer product sectors, e.g., shoes, pharmaceuticals.

Tom Vest, Senior Policy Advisor, and KC Claffy, Principal Investigator, Cooperative Association for Internet Data Analysis (CAIDA).

Addendum: "If you Can't Measure It, You Can't Manage It & if you can't manage it - you can't govern it", by Bill St Arnaud, CANARIE

All:

I would like to add an addendum to kc and Tom Vest’s phrase “If you can’t measure it, you can’t manage it” to the effect that “If you can’t manage it, you can’t govern it”

Many people still think of the Internet as a homogenous network architecture made up of router, IP packets and physical links. But, in fact we are already seeing increased specialization with many “Internet” networks dedicated to specific applications and/or services. For example many organizations are building dedicated VoIP networks and VoIP peering and exchange points. These are lot different than traditional Internet peering and exchange points. As well many companies such as IBM, EDS, etc are building networks dedicated to web service interaction and grid services. In addition there are many overlay P2P networks with their own independent naming and routing systems. The university research community is also experimenting with private Internet networks dedicated to specific uses or applications e.g Teragrid, CA*net 4 user controlled lightpaths, EGEE, etc etc. The advent of “infrastructure-free” applications like Skype, delicious etc will further compound the complexity of the Internet.

The bottom line is that increasingly the Internet is becoming far more complex and specialized and as kc has discovered increasingly difficult, if not impossible to measure from any single central vantage point such as major public peering facility. Private networks, application specific peering, etc are only going to make this more difficult in the future. I also agree with kc that if you can’t measure a network a network, you can’t manage it – at least from a central carrier vantage point. The only place where you will be able to measure traffic is at the edge for a given user or customer. But getting overall measurement of the Internet traffic will be impossible.

In many ways this reflects the challenges economists face. They can get very detailed micro-economic statistics for a particular company or industry and do detailed In/out analysis, but they must rely on statistical sampling and aggregate data (which is fraught with errors and misinterpretation) for macro-economic analysis.

To my mind the Internet is evolving in complexity like the economy itself. In many ways the Internet of today is equivalent to the simple barter system of primitive economic society in terms of technology maturity. If the Internet is a general purpose technology (GPT) that will be an integral part of our economy and society we need to work with economists to develop new statistical and aggregate data tools to measure the aggregate impact of the Internet. This is why I applaud Tom Vest’s seminal work in this area.

But the more profound implication of the increased complexity of the Internet is that just as it is impossible for governments to mico-manage the economy (but this does stop them from trying) I think it will be just as equally difficult to establish any type of micro-management governance structure for the Internet (but as with government it will not stop various stakeholders and those with vested interested from trying to do so).

Bill

A Day in the Life

In 2001 the U.S. National Academy of Sciences convened a workshop to assess the state of networking research, and, in pursuit of objectivity and fresh insights, arranged for more than half of the attendees to be from other fields, in this case computer science. Among the most memorable conclusions:

.. the outsiders expressed the view that the network research community should not devote all or even the majority of its time to fixing current Internet problems.

Instead, networking research should more aggressively seek to develop new ideas and approaches. A program that does this would be centered on the three M's -- measurement of the Internet, modeling of the Internet, and making disruptive prototypes. These elements can be summarized as follows:

Measuring -- The Internet lacks the means to perform comprehensive measurement on activity in the network. Better information on the network would provide the basis for uncovering trends, as a baseline for understanding the implications of introducing new ideas into the network, and would help drive simulations that could be used for designing new architectures and protocols. This report challenges the research community to develop the means to capture a day in the life of the Internet to provide such information.

Modeling -- The community lacks an adequate theoretical basis for understanding many pressing problems such as network robustness and manageability. A more fundamental understanding of these important problems requires new theoretical foundations -- ways of reasoning about these problems that are rooted in realistic assumptions. Also, advances are needed if we are to successfully model the full range of behaviors displayed in real-life, large-scale networks.

Making disruptive prototypes-- To encourage thinking that is unconstrained by the current Internet, Plan B approaches should be pursued that begin with a clean slate and only later (if warranted) consider migration from current technology. A number of disruptive design ideas and an implementation strategy for testing them are described in Chapter 4.

-- National Academies Press, "Looking over the Fence at Networks: A Neighbor's View of Networking Research (2001)"

Per the above -- "This report challenges the research community to develop the means to capture a day in the life of the Internet" -- we admit that the research community has not come anywhere near this goal, nor does it seem a priority. We seek to open a discussion on what it would mean, require, and cost to capture a day in the life of the Internet with as much scientifically grounded methodology as possible, and with resulting data as widely accessible as possible. We recognize that the proposed project will involve building a cooperative community to support the simultaneous capture of a variety of measurements from and across many strategic links around the globe for further analysis by research scientists. But by establishing a periodic tradition of synchronized measurements, and supporting tools, analysis, visualization, and data catalog ( DatCat, http://www.datcat.org, Internet Traffic Archive, CRAWDAD, MOME, Datapository, PREDICT) in which to index collected traces, we hope to significantly increase the quantity as well as quality of empirical data supporting Internet research.

Several complementary projects at CAIDA provide the impetus for our first attempt to coordinate a distributed measurement activity in late 2006. As part of an NSF-sponsored DNS measurement project ( http://www.caida.org/funding/dns-itr/ ), CAIDA and ISC plan to perform a 48-hour simultaneous measurement event on dozens of root server anycast nodes. Specifically, ISC will collect packet header traces from multiple (hopefully all) anycast instances of at least three root nameservers, based on feedback from the previous such measurement experiment. ( DNS measurement Recommendations ) Since to our knowledge this event will be the largest scale simultaneous collection from a core component of the global Internet infrastructure, we consider it an ideal time to prototype a "Day in the Life of the Internet" measurement event. Specifically, if you have access to or influence over Internet measurement infrastructure and can contribute datasets (anonymized according to your needs), please email ditl-info@caida.org for details regarding already planned measurement dates, times, locations, and types of data. (There will be an informal vetting process to avoid manipulation of the experiment.)

We also seek input from others interested in gathering specific complementary measurements on the same days, to help us maximize the return on investment of participation in the experiment.

Commercial pressures make it next to impossible to get Internet measurement data to the research community, but empirical network science is not possible without such data. We hope that over time, annual measurement activities to support "day in the life of the Internet" (DITL) data sets will gather increasing momentum. Ideally, participating partners would provide simultaneous capture of a variety of trace data: workload, topology, routing, and performance, from a large number of strategic locations around the globe, anonymized appropriately according to local restrictions.

We recognize this project will involve global efforts to overcome logistic, technical, economic, and legal obstacles to measurement and data sharing. But through this activity we seek to determine whether, given enough interest in the community, it is possible to gather sufficient data not only to capture salient characteristics of `a day in the life of the Internet', but also to provide sufficient empirical grounding for the development of reliable predictive models of Internet traffic, topology, routing, and evolution.