Skip to main content

According to the Best Available Data

The CAIDA blog — commentary and analysis of Internet measurement, infrastructure, policy, and economics, published since 2006. Formerly hosted at blog.caida.org. See About the CAIDA Blog · Categories · All entries by month · RSS feed.

IPv4 and IPv6 AS Core 2013

We recently released a visualization at http://www.caida.org/research/topology/as_core_network/ that represents our macroscopic snapshots of IPv4 and IPv6 Internet topology samples captured in 2013. The plots illustrate both the extensive geographical scope as well as rich interconnectivity of nodes participating in the global Internet routing system.

IPv4 and IPv6 AS Core Graph, Jan 2013

This AS core visualization addresses one of CAIDA's topology mapping project goals is to develop techniques to illustrate structural relationships and depict critical components of the Internet infrastructure. These IPv4 and IPv6 graphs show the relative growth of the two Internet topologies, and in particular the steady continued growth of the IPv6 topology. Although both IPv4 and IPv6 topologies experienced a lot of churn, the net change in number of ASes was 3,290 (10.7%) in our IPv4 graph and 495 (25.7%) in our IPv6 graph.

In order to improve our AS Core visualization over previous years, this year we made two major refinements to our graphing methodology, including how we rank individual ASes. First, we now rank ASes based on their transit degree rather then their outdegree. Second, we now infer links across Internet eXchange (IX) point address space, rather than considering the IX itself a node to which various ISPs attach. Details at http://www.caida.org/research/topology/as_core_network/.

[For details on a more sophisticated methodology for ranking AS interconnectivity, based on inferring AS relationships from BGP data, see http://www.caida.org/data/active/as-relationships/.]

CAIDA's Annual Report for 2012

[Executive Summary from our annual report for 2012.]

This annual report covers CAIDA's activities in 2012, summarizing highlights from our research, infrastructure, data-sharing and outreach activities. Our research projects span Internet topology, routing, traffic, economics, future Internet architectures, and policy. Our infrastructure activities continue to support measurement-based studies of the Internet's core infrastructure, with focus on the health and integrity of the global Internet's topology, routing, addressing, and naming systems. In 2012 we increased our participation in future Internet research in two dimensions: measuring and modeling IPv6 deployment; and an expanded role (in management) of the Named Data Networking project, one of the NSF-funded future Internet architecture projects headed into its third year. We also began a project to study large-scale Internet outages via correlation of a variety of disparate sources of data.

We continued to make advances in Internet topology research, supported by our expanding Ark measurement infrastructure. We collect and share the largest Internet topology data sets (IPv4 and IPv6) available to academic researchers, and we share many aggregated annotated derivative data sets publicly, including rankings of ISPs annotated with (our estimated) business relationships between autonomous networks. Our topology measurement platform supports IPv6 -- by the end of 2012, 28 of our 64 Ark hosting sites provided IPv6 connectivity and topology measurements. Using our new alias resolution measurement system, which integrates and improves on the best available technology for IP address alias resolution, we collected, analyzed, processed and released our fifth published Internet Topology Data Kit (ITDK), reflecting measurements taken in July 2012. The July 2012 ITDK includes two related router-level topologies, router-to-AS assignments; geographic location of each router; and DNS lookups of all observed IP addresses. After an extensive exercise with our validation data via AS Rank, we also spent many months this year overhauling our AS relationship inference algorithm so that we can add AS relationship annotations to future ITDKs.

On the theoretical side of topology research, we developed a new model in which new connections optimize certain trade-offs between popularity and similarity of nodes, instead of simply preferring popular nodes. This framework has a geometric interpretation in which popularity preference emerges from local optimization. In contrast to standard preferential attachment, our optimization framework accurately describes the large-scale evolution of technological (the Internet), social (trust relationships between people) and biological (Escherichia coli metabolic) networks, accurately predicting the probability of new links. We developed a related framework to support mapping a real network into a hyperbolic plane in a way congruent with this model of network growth. Perhaps our most exciting theoretical result was our discovery of structural similarity (power-law graph with strong clustering) between a casual network representing the large-scale structure of spacetime in our accelerating universe, and complex networks such as the Internet, social, or biological networks. We collaborated with supercomputing experts at SDSC to run HPC simulations that provided evidence that this structural similarity is due to asymptotic equivalence in large-scale growth dynamics of complex networks and spacetime in the universe.

In 2012 we continued applying our theoretical, empirical, and practical understandings of the Internet's evolution to the challenge of enable dramatically more scalable global Internet routing. We continued our partnership in the Named Data Networking project, a 12-university collaboration funded by NSF's Future Internet Architecture (FIA) Research program to explore a generalization of the Internet architecture that allows naming more than just communication endpoints, i.e, the source and destination IP addresses, but also data (content) itself. This approach shifts the focus from where -- addresses and hosts in today's Internet -- to what -- the content that users and applications care about. By naming data instead of locations, the new architecture transforms data into a first-class entity while addressing the known technical challenges of today's Internet: routing scalability, network security, content protection and privacy. In 2012 we investigated combinations of name-space structure and network topology that optimize the efficiency of NDN algorithms and participated in NDN testbed development and evaluation. The most challenging part of this routing research as it pertains to the Internet still lies ahead, and will require a broader community of engaged thinkers: application of these and other theoretical results to real-world Internet security, economic, and policy contexts.

A more immediate architectural need of the global Internet has inspired us to study the transition to IPv6. The two main lessons we can glean from the scant data available are: (i) architectural transitions - even those deemed minor but essential - are slow; (ii) the U.S. is behind other regions of the world in IPv6 deployment, and has not thus far invested in shedding quantitative light on this problem, despite making attempts to lightly nudge the market toward wider IPv6 adoption. With support from NSF, we collaborated with the Naval Postgraduate School (Rob Beverly) we studied the deployment of IPv6 at the Autonomous System (AS) level using historical BGP data and recent active measurements, to compare IPv4 topology structure and adoption trends. While most core Internet transit providers have deployed IPv6, edge networks are lagging. IPv6 deployment is stronger in Europe and the Asia-Pacific region, than in North America. The IPv6 topology is characterized by a single dominant player, Hurricane Electric, which appears in a large fraction of IPv6 AS paths, and is more dominant in IPv6 than the most dominant player in IPv4. Routing dynamics in the IPv6 topology are largely similar to those in IPv4, and churn in both networks grows at the same rate as the underlying topologies. We found that performance over IPv6 paths is comparable to that over IPv4 paths if the AS-level paths are the same, but can be much worse than IPv4 if the AS-level paths differ. To support a separate but related modeling effort, we developed and conducted a survey of network operators to gauge IPv6 deployment patterns and plans. Based on the results we hope to refine and re-issue a survey next year, to inform and parameterize a predictive model of possible IPv6 future trajectories.

We made significant progress on our Internet economics research, one goal of which is to create a scientific basis for modeling Internet interdomain interconnection and dynamics, capturing relevant interactions between network business relations, internetwork topology, routing policies, and resulting interdomain traffic flow. We developed and published a holistic cost model that can help operators evaluate the costs of various routing and peering decisions, among other network operation costs. Using traffic data from a large carrier network, our model revealed how network operators can significantly reduce the cost of carrying traffic in their networks by adjusting routing for a small fraction of total traffic. We also published a paper on our GENESIS simulator, which embodies a computational model of interdomain network formation that captures key factors influencing network formation dynamics: highly skewed traffic matrix, policy-based routing, geographic co-location constraints, and the costs of transit/peering agreements. This simulator enables us to study ``what-if'' questions, such as asking how open peering strategies affect networks in terms of topology, traffic flow, and financial health. We continued studying available interdomain traffic matrix (ITM) data, and discovered that we can model the traffic sent by an AS as either a log-normal or Pareto distribution, depending on whether congestion levels. We found correlations between different ASes mostly due to relatively few highly popular prefixes. We also held a successful interdisciplinary Workshop on Internet Economics (WIE) in December 2012 (co-hosted with MIT's Dave Clark), focused on reaching consensus on definitions and data to support a regulatory framework for a converged communications infrastructure.

In early 2012, we undertook a new three-year research effort to study large-scale Internet outages, under an exciting new Transition to Practice area of NSF's Secure and Trustworthy Cyberspace research program. In this project we are applying our successful results in studying the Egypt and Libya censorship-induced outages (our IMC2011 paper) to the development, testing, and deployment of an operational capability to detect, monitor, and characterize future episodes of Internet connectivity disruptions. In early 2012, we published a study that used the UCSD darknet traffic data to analyze other outages caused by geophysical disasters -- the earthquakes in Christchurch and Tohoku in 2011 -- which won an ACM SIGCOMM CCR award for one of the best CCR papers of 2012.

We continued to dedicate resources to support the infrastructure measurement and data sharing interests and needs of two U.S. federal agency programs: the National Science Foundation's International Research Network Connections (IRNC) program, and the Department of Homeland Security's Protected Repository of Data on Internet CyberThreats (PREDICT) data-sharing project (http://www.predict.org). The PREDICT funding provides essential support for deployment and operations of our measurement infrastructure, and the collection, curation, and sharing of several unprecedented data sets available to researchers (http://www.caida.org/data/). We are responsive to researcher requests for additional/different Internet data sets, to the extent possible given our resources. We have found an increasing number of disciplines (physicists, sociologists, biologists) interested in our Internet measurement data sets and research results as they apply to other complex network structure, behavior, and evolution.

Finally, as always, we engaged in a variety of tool development, data-sharing, and outreach activities, including web sites, 16 peer-reviewed papers, 5 technical and workshop reports, 47 presentations, 13 blog entries, 7 animations, and (six) workshops, and a seminar series.

Full annual report:
http://www.caida.org/home/about/annualreports/2012/

Program plan for 2010-2013:
http://www.caida.org/home/about/progplan/progplan2010/

We will be creating a new 3-year program plan in 2013. Please do not hesitate to send comments or questions to info at caida dot org.

network mapping and measurement conference

I had the honor of presenting an overview of CAIDA's recent research activities at the Network Mapping and Measurement Conference hosted by Sean Warnick and Daniel Zappala. Talks topics included: social learning behavior in complex networks, re-routing based on expected network outages along current paths, twitter data mining to analyze suicide risk factors and political sentiments (three different talks). James Allen Evans gave a sociology of science talk, an interview form of which seems to be achived by the Oxford Internet Institute. The organizers even arranged a talk from a local startup, NUVI, doing some fascinating real-time visualization and analytics of social network data (including Twitter, Facebook, Reddit, Youtube).

The workshop was held at Sundance, Utah, one of the most beautiful places I've ever been for a workshop. This workshop series was originally DoD-sponsored with lots of government attendees interested in Internet infrastructure protection, but sequester and travel freezes this year yielded only two USG attendees, and budget constraints may keep this workshop from happening again next year. I hope not, it was really a unique environment and exposed me to a range of work I would not otherwise have discovered anytime soon. Kudos to the organizers and sponsors.

Carna botnet scans confirmed

On March 17, 2013, the authors of an anonymous email to the "Full Disclosure" mailing list announced that last year they conducted a full probing of the entire IPv4 Internet. They claimed they used a botnet (named "carna" botnet) created by infecting machines vulnerable due to use of default login/password pairs (e.g., admin/admin). The botnet instructed each of these machines to execute a portion of the scan and then transfer the results to a central server. The authors also published a detailed description of how they operated, along with 9TB of raw logs of the scanning activity.

Online magazines and newspapers reported the news, which triggered some debate in the research community about the ethical implications of using such data for research purposes. A more fundamental question received less attention: since the authors went out of their way to remain anonymous, and the only data available about this event is the data they provide, how do we know this scan actually happened? If it did, how do we know that the resulting data is correct?

Since we could not find any third-party validation of this event, we looked for evidence in the traffic captured at the UCSD Network Telescope (a large darknet). From this traffic we selected probing packets consistent with the default nmap host probe (comprised of four different types of packets) that the carna botnet used. The visualization below shows, for each day of 2012, the total number of probes we observed at the telescope in bins of 1 day (blue line). While these probes may have been generated by any host on the Internet, the large increase visible between April and September 2012 matches the logs distributed by the authors of the botnet (red line), showing evidence of this scanning activity.


We also found that some of the raw logs of the carna botnet erroneously reported that a large number of IPs in our darknet were active, and specifically accepting connections on port TCP 80 (darknet IP addresses are inactive by definition, thus not accepting connections). A preliminary analysis suggests that this measurement error is likely due to the presence of HTTP proxies in some of the networks that hosted scanning bots. The default nmap host probe sends four different packets trying to solicit a response from the target: (i) ICMP echo request, (ii) ICMP timestamp, (iii) TCP ack on port 80, (iv) TCP syn on port 443.  For darknet addresses that the carna logs report as inactive, we observed all four of these packets, but for the addresses misreported as active, packets of type (iii) did not reach the telescope. We suspect that these packets were intercepted by HTTP proxies whose replies caused the bots to falsely report the target IP address as listening on port TCP 80.
Assuming these bots probed the rest of the IPv4 Internet proportionally to their probing of the darknet we can observe, about 3% of the host probe logs and port scan logs of the carna botnet could potentially be affected by this particular problem. The maps and animations they published seem unaffected by this issue because they were based on ICMP pings and actual (application-layer) responses from the target hosts.

We have only briefly investigated the carna botnet scan, but there are clearly epistemological issues related to any potential scientific use of the data published by the botnet authors. There are even more complex ethical issues related to using this data set, as well as with its original collection. We have previously mentioned efforts to provide ethical guidance to Internet researchers; the debate continues and this data set will likely become an interesting part of it.

Third Workshop on Internet Economics (WIE2012)

As part of our NSF-funded network research project on modeling Internet interconnection dynamics, David Clark (MIT) and I hosted the second Workshop on Internet Economics (WIE2012) last December 12-13. The goal of the workshop was to provide a forum for researchers, commercial Internet facilities and service providers, technologists, economists, theorists, policy makers, and other stakeholders to empirically inform emerging regulatory and policy debates. The theme for this year’s workshop was "Definitions and Data". The final report describes the discussions and presents relevant open research questions identified by workshop participants. Slides presented at the workshop are available at the workshop home page. From the intro (but the full report (6-page pdf) is worth reading):

Building on the success of our first two workshops in this series [WIE09,WIE11], we held the 3rd Workshop on Internet Economics (WIE). The theme for this year's workshop was "Definitions and Data", motivated by our sense that many of the debates today about effective regulation are clouded by lack of clarity about terms and concepts, and lack of real information about the current state of the communications infrastructure. Concepts that have resisted clean definition include network neutrality, reasonable network management, market power, and reliability. Stakeholders disagree on fundamental parameters of central concepts in the industry, such as interconnection, or the metrics for broadband quality itself.

Equally missing is good data on what is actually happening. Whether measurements are undertaken by the FCC, as with the current SamKnows effort, or by the research community or industry, good definition must precede good measurement, because collectively we must be consistent and clear what we are proposing to measure and why. A guiding premise of this workshop was that attention to definitions can inform research in data gathering, which in turn can inform regulatory debate. Workshop discussions also focused on the impacts of the limitations of currently available data (such as undersampling) and how to gain more relevant data with minimal impact on personal privacy.

The workshop format focused discussion around six pre-selected topics: defining broadband (wired and wireless); Interconnection; definitions and metrics of market power; the emergence of private IP networks; regulatory distinctions in a converged world; and defining acceptable practice for data-gathering. We spent about two hours per topic, with at least two 10-minute talks followed by an hour for each discussion. Three promising future research directions emerged. First, we reached rough consensus on a proposed practical approach to measure a user's ``quality of experience'' (QoE), one that could frame not only a stable definition of broadband Internet service but also enable more rigorous description of ``willingness to pay'' for different applications. Second, most participants agreed that the rise of private IP networks as an alternative platform to the public Internet (and to the economically unsustainable PSTN) promise an even more opaque future at a time when it has become clear that much of current communications regulation lacks empirical basis. Third, there was recognition that both scientific research and sound public policy share the need to develop, maintain, and archive some classic data sets to develop some sense of history and to inform general models of network behavior. One possible goal for a future workshop is try articulate an argument for data that might be valuable in the future, not only to support specific policy questions but also to begin to establish historical baselines and promote scientific inquiry. More formal and transparent ties between policymakers and researchers could frame ethical use of such data.

Subsequent sections of the report are:
2. Regulatory Distinctions Amid Convergence
3. Defining Broadband
4. Defining Market Power
5. Interconnection
6. The Emergence of Private IP Networks
7. Acceptable Practices for Data-gathering
8. Future Research Directions

Full report: http://www.caida.org/publications/papers/2013/wie2012_report/ Other materials from workshop: http://www.caida.org/workshops/wie/1212. Feedback welcome. Thanks to all who participated.

Correlation between country governance regimes and the reputation of their Internet (IP) address allocations

[While getting our feet wet with D3 (what a wonderful tool!), we finally tried this analysis tidbit that's been on our list for a while.]

We recently analyzed the reputation of a country's Internet (IPv4) addresses by examining the number of blacklisted IPv4 addresses that geolocate to a given country. We compared this indicator with two qualitative measures of each country's governance. We hypothesized that countries with more transparent, democratic governmental institutions would harbor a smaller fraction of misbehaving (blacklisted) hosts. The available data confirms this hypothesis. A similar correlation exists between perceived corruption and fraction of blacklisted IP addresses.

For more details of data sources and analysis, see:
http://www.caida.org/research/policy/country-level-ip-reputation/

x:Corruption Perceptions Index
y:IP population %
x:Democracy Index
y:IP population %
x:Democracy Index
y:IP infection %

Interactive graph and analysis on the CAIDA website

2001:deba:7ab1:e::effe:c75

[This blog entry is guest written by Robert Beverly at the Naval Postgraduate School.]

In many respects, the deployment, adoption, use, and performance of IPv6 has received more recent attention than IPv4. Certainly the longitudinal measurement of IPv6, from its infancy to the exhaustion of ICANN v4 space to native 1% penetration (as observed by Google), is more complete than IPv4. Indeed, there are many vested parties in (either the success or failure) of IPv6, and numerous IPv6 measurement efforts afoot.

Researchers from Akamai, CAIDA, ICSI, NPS, and MIT met in early January, 2013 to firstly share and make sense of current measurement initiatives, while secondly plotting a path forward for the community in measuring IPv6. A specific objective of the meeting was to understand which aspects of IPv6 measurement are "done" (in the sense that there exists a sound methodology, even if measurement should continue), and which IPv6 questions/measurements remain open research problems. The meeting agenda and presentation slides are archived online.

To this end, it's important to note that one of the central observations of claffy's CCR editorial from July 2011 is that the eventual fate of IPv6 remains undecided. Similarly, whether there will be a "forcing function" that leads to non-trivial IPv6 deployment is still unclear.

Thomas Blood of NPS's central IT organization spoke to this point with respect to DoD IPv6 compliance mandates where all internal DoD networks should be IPv6 compliant by 2014. The deadline for universal DoD v6 compliance has passed a remarkable four times already, with only 1% of organizations meeting the mandate as of June 2012. Despite pressure to adopt IPv6 within the US government, funding and security issues trump most efforts to deploy IPv6. Without any demand from users, there is little incentive to deploy IPv6 -- especially given IT personnel effort in supporting and securing v6. There is not only a large amount of legacy equipment that cannot support IPv6, but also a wide range of products on the market today that do not properly support IPv6. Indeed, DREN has performed extensive vendor IPv6 testing, with mixed results, for example, discovering essential devices that do not support IPv6 ACLs, or those that do so with unacceptably low performance.

It is important for existing measurement efforts and methodologies to continue to collect data during this potential evolution -- to understand the adoption of IPv6, or to better understand why IPv6 fails if its adoption languishes. Several components of the "data we need" have largely been addressed in the last year. For instance, while methodologies to assess IPv6 adoption are now well-understood, continued data collection is important. On the client-side, recent work from Zander et al. at IMC 2012 showed a clever way to leverage Google's vast visibility of the edge to obtain a large and diverse sample of client IPv6 capability by embedding measurements into flash advertisements. Google and Akamai both publish IPv6 adoption data based on observing client behavior. Google now publishes non-whitelisted AAAA DNS records for its domains, and supports IPv6 end-to-end. On the server-side, significant insight can be gleaned from the DNS. For example, ICSI, in collaboration with the University of Michigan, is analyzing zone and query data from some of the TLD authorities to understand the penetration of infrastructure IPv6. While Claffy's 2011 editorial noted that estimates of IPv6 penetration vary by orders of magnitude, these measurements are slowly converging to a more reliable estimate.

Performance comparisons between IPv4 and IPv6 have also received significant attention, including CAIDA's IMC 2012 paper and Akamai's measurements. Several independent sources have shown that, when paths are congruent at the AS-level, performance is largely the same -- i.e. there is no data-plane performance penalty today to using IPv6. It will be important to continue performance measurements as more applications implement happy eyeballs.

Workshop participants also identified areas where additional research is needed: a) measuring the extent of carrier grade NAT in the Internet; b) understanding IPv6 topology; c) characterizing IPv6 security issues; d) incorporating economic models.

Given IPv4 address exhaustion and economic incentives against adopting IPv6, providers may choose to deploy carrier grade NAT rather than (or in addition to) investing in IPv6. Little data exists today to understand the extent and use of carrier grade NAT. Arthur Berger and Nick Weaver exchanged several ideas for finding carrier grade NATs during the meeting and Nick hopes to place new functionality into Netalyzer to more broadly study their deployment.

With respect to economics, Steven Bauer presented a thought-provoking assessment of the (lack of) time-series IPv6 data in residential broadband performance studies (e.g., Samknows, Bismark, etc). In particular, when evaluating speed measurements and overall user experience, how should one weight IPv4 versus IPv6? Further, recent work that has evaluated the graph-theoretic centrality of AS interconnection in an effort to quantify market power (e.g. with respect to recent de-peering arguments, etc), have not examined IPv6 peering -- suggesting avenues of exploration that may be valuable in understanding IPv6.

Topology is a second area where slow but steady progress is being made -- yet much more remains to be done. Researchers from NPS and CAIDA are collaborating on efforts to perform IPv6 alias resolution to reduce interface-level topologies (i.e. as collected by traceroute) to the more useful router-level topologies. Their most recent work will appear at PAM 2013; they are continuing the collaboration with a focus on scaling to Internet-size topologies, i.e., large-scale IPv6 alias resolution. Further, researchers from NPS, Akamai, and ICSI are collaborating on efforts to infer "sibling" relationships between IPv4 and IPv6 addresses, with the eventual goal of enabling comparative topology mapping, sound performance comparisons, and informing reputation and geolocation engines.

Lastly, there was consensus among participants that security is one of IPv6's Achilles' heels. At the meeting, Chris Eagle, one of the organizers for previous Defcon CTF exercises, spoke about their recent use of IPv6 in the challenge, which pretty much stumped all the security experts participating in the contest. IPv6 as deployed today is largely unsecured with respect to known vulnerabilities in IPv4, while introducing a raft of new attack vectors. As a result, NPS is undertaking an effort to better support IPv6 in the spoofer project, as well as continuing to use IPv6 attack traffic for opportunistic measurement insight.

Much exciting work is in progress!

Packet Loss Metrics from Darknet Traffic

At the CoNEXT Student Workshop, in Nice, France on December 10, 2012, CAIDA shared recent research on Internet outages in a poster entitled "Gaining Insight Into AS-Level Outages through Analysis of Internet Background Radiation."

An initial task of our NSF-funded DALS SaTC project is to refine and extend indicators to support real-time detection and rapid characterization of Internet connectivity outage events. We used several darknet-based metrics in our studies of country-wide censorship and the impact of political and geophysical events. This latest metric characterizes the number of TCP SYN packets sent from selected networks to the UCSD Network Telescope, to help determine whether packet loss (e.g., because of congestion) is associated with the outage.  Since the UCSD Network Telescope receives traffic sent to unassigned IP addresses but does not respond, TCP connection attempts are comprised of only SYN packets. Conficker-like packets comprise the vast majority of these packets, known as Internet Background Radiation.  Conficker-infected hosts are known to send two SYN packets per connection attempt, a consistent behavior that allows us to infer packet loss when the number of packets per connection attempt decreases for this type of traffic.

The poster highlights two case studies. In the "Dodo-Telstra" Routing Leakage, caused by a BGP leak, the metric γ decreases significantly, consistent with a bottleneck preceding the outage.

However, during the Libyan Internet Blackout of 2011, where the Libyan government used packet filtering to implement country wide censorship, the value of  γ did not change when a few hosts were allowed through the "firewall".  This behavior is consistent with filtering decreasing the number of sources sending traffic without changing its per-flow characteristics.

Our poster was voted as one of the top 8 of the student workshop; these 8 were presented at the main CoNEXT conference.

Syria disappears from the Internet

On the 29th of November, shortly after 10am UTC (12pm Damascus time), the Syrian state telecom (AS29386) withdrew the majority of BGP routes to Syrian networks (see reports from Renesys, Arbor, CloudFlare, BGPmon). Five prefixes allocated to Syrian organizations remained reachable for another several hours, served by Tata Communications. By midnight UTC on the 29th, as reported by BGPmon, these five prefixes had also been withdrawn from the global routing table, completing the disconnection of Syria from the rest of the Internet.

Several organizations with access to different sources of data that illuminate aspects of the blackout have released their data analyses. Renesys and BGPmon used BGP routing data to monitor the systematic withdrawal of routes to networks in Syria. Arbor Networks used traffic flow data collected from their globally distributed ATLAS infrastructure, which serves hundreds of customers. Akamai has traffic data from their own content distribution network infrastructure, and released a graph showing an abrupt drop in the volume of (HTTP) traffic Akamai servers sent to Syrian hosts. While the RIPE NCC allowed users to follow the BGP update activity for Syrian prefixes in near-realtime.

We provide another lens through which the blackout could be observed: a drop in unsolicited traffic generated by malware-infected Syrian PCs. Malware (worms, viruses, etc) often spreads to other vulnerable computers over the Internet by way of random scanning by infected hosts. A signal-producing side effect of a country-level Internet blackout is that Internet access is also denied to malware attempting to infect other hosts. This drop in unsolicited traffic can be observed in data captured from a darknet such as the UCSD Network Telescope. A darknet is a block of globally reachable but unassigned IP addresses; all traffic destined to such addresses is unsolicited, most of it from malware-infected PCs. We have previously used this technique to analyze the Internet blackouts in Egypt and Libya during the Arab Spring uprisings of last year and the impact of the earthquakes in Japan and New Zealand in early 2011.

The Syrian Internet Blackout in Nov 2012 as seen at the UCSD Network Telescope
The Syrian Internet Blackout in Nov 2012 as seen at the UCSD Network Telescope

This graph shows the number of unique Syrian source IP addresses per hour sending traffic that reaches the UCSD Network Telescope. Our data confirms the findings of other groups, showing an abrupt decrease in the number of transmitting Syrian hosts between 10 and 11am UTC on the 29th. For the following 48 hours we received almost no traffic from Syrian hosts. To determine that an IP address belongs to a Syrian host, we constructed a list of prefixes officially delegated by RIPE NCC to Syrian organizations, augmented with the 5 prefixes advertised by Tata Communications (as reported by BGPmon), which were the last to be withdrawn. We then validated the addresses found in the telescope data against the Maxmind GeoLite Country database and through manual traceroutes.

During the period of the blackout we received a total of 6 packets from 3 sources inside Syrian address space. These packets had source IP addresses (which could be spoofed, we are still investigating) within the networks advertised by Tata Communications. We observed this traffic after these routes had been withdrawn (according to BGPmon), so it is possible that some Syrian networks were still able to send traffic by way of default routes, as was the case for some hosts during the Egyptian blackout (see our IMC2011 paper). Traffic began returning to pre-blackout levels just after 2pm UTC on December 1st.

This activity is part of our NSF SATC-funded project on Internet outages (NSF CNS-1228994), and is also supported by measurement and data curation made possible by DHS S&T's PREDICT and Cybersecurity programs (Cooperative Agreement FA8750-12-2-0326 and Contract N66001-12-C-0130).

Team: Alistair King, Karyn Benson, Brad Huffaker, Marina Fomenkov, Emile Aben, Alberto Dainotti, KC Claffy

CAIDA at the NSF Secure and Trustworthy Cyberspace (SaTC) Principal Investigators' Meeting

Last week CAIDA researchers (Alberto and kc) visited National Harbor (Maryland) for the 1st NSF Secure and Trustworthy Cyberspace (SaTC) Principal Investigators Meeting. The National Science Foundation's SATC program is an interdisciplinary expansion of the old Trustworthy Computing program sponsored by CISE, extended to include the SBE, MPS, and EHR directorates. The SATC program also includes a bold new Transition to Practice category of project funding -- to address the challenge of moving from research to capability -- which we are excited and honored to be a part of.

This PI meeting included social science, economic, policy, as well as technical perspectives on cybersecurity through plenary talks, breakout sessions, posters, and an adventurous one-on-one researcher "speed dating" experiment. We presented a poster that summarized our current NSF SATC-funded effort to build a platform for online monitoring and analysis of large-scale Internet infrastructure outages. The poster (reproduced below) displays highlights of our previous results from analyzing large outages in Egypt and Libya during the so called "Arab Spring", and the impact of the earthquakes in Japan and New Zealand in 2011. More soon, as we are still analyzing the most recent large-scale Internet outage in the news (Syria).