Monday, October 19, 2009

PART 2: TELEPRESENCE INTEROPERABILITY CHALLENGES

The challenges around telepresence interoperability are related to both logistics and technology. Logistics are probably the bigger problem. Vendors usually conduct interoperability tests by gathering at an interoperability event, bringing their equipment to a meeting location, and running test plans with each other. This is the way IMTC manages H.323 interoperability tests and also the way SIPit http://www.sipit.net/ manages tests based on the SIP protocol. While developers can pack their new video codec in a suitcase and travel to the meeting site, multi-codec telepresence systems are large and difficult to transport. A full-blown telepresence system comes on a large truck and takes substantial time to build – usually a day or more. Therefore, bringing telepresence systems to interoperability test events is out of the question.

An alternative way to test interoperability is by vendors purchasing each other’s equipment and running tests in their own labs. While this is acceptable approach for $10K video codecs, it is difficult to replicate with telepresence systems that cost upwards of $200K. One could ask “Why don’t you just connect the different systems through the Internet for tests?” The issue is that telepresence systems today are run on fairly isolated segments of the IP network - mostly to guarantee quality but also due to security concerns - and connecting these systems to the Internet is not trivial. It requires rerouting network traffic, and use of video border proxies to traverse firewalls.

The technology challenges require more detailed explanation. Vendors like HP and Teleris run closed proprietary telepresence networks and their telepresence systems cannot talk directly to other vendor’s systems. There are of course gateways that can be used for external connectivity but gateways mean transcoding, i.e. decrease of quality, limited capacity, and decreased reliability of end-to-end communication. For those not familiar with the term ‘transcoding’, it is basically translation from one video format into another video format. Telepresence systems send and receive HD video at 2-10 megabits per second (Mbps) for each screen/codec in the system, and all that information has to go through the gateway and be translated into a format that standards-based systems can understand.

Some telepresence vendors state that they support standards such as H.323 or SIP (Session Initiation Protocol). However, standard- compliance is not black-and-white, and telepresence systems can support standards and still not allow good interoperability with other vendors’ systems. When Cisco introduced its three-screen CTS 3000, they made the primary video codec multiplex three video streams – its own and the two captured by the other two codecs – into a single stream that traversed the IP network to the destination’s primary codec. Third-party codecs cannot understand the multiplexed bit stream, and that is basically why you cannot connect a Polycom, LifeSize, or Tandberg telepresence system to Cisco CTS. Note that Cisco uses SIP for signaling and claims therefore standard-compliance; however, the net result is that third-party systems cannot connect. If you decide to spend more money and buy a gateway from Cisco, you could connect to third-party system but at a decreased video and audio quality that is far from the telepresence promise of immersive communication and replacement of face-to-face meetings. The discrepancy between the ‘standard compliance’ claim and the reality that its systems just do not talk to any other vendor has haunted Cisco since they entered the video market.

When Tandberg introduced its three-screen telepresence system T3, they made another technological decision that impacts interoperability. T3 combines the video streams from three codecs (one per screen) into one stream, and any non-Tandberg system that connects to T3 receives what we call a ‘filmstrip’, i.e. three small images next to each other (http://www.flickr.com/photos/20518315@N00/4015164378/). The ‘filmstrip’ covers maybe one-third of one screen (or one-ninth of the total screen real estate of a three-screen system). So, yes, you can connect to T3 but you lose the immersive, face to face feeling that is expected of a telepresence system. Note that T3 uses standard H.323 signaling to communicate with other systems, so it is standard-compliant; however, the result is that if you want to see the three images from T3 on full screens, you have to add an expensive Tandberg Telepresence Server (TTPS). I will discuss TTPS in more detail in parts 4 and 5.

To come back to my original point, due to a range of logistical and technological issues, establishing telepresence interoperability is quite a feat that requires serious vendor commitment and a lot of work across the industry.

Stay tuned for Part 3 about the organizational issues around telepresence interoperability testing … http://videonetworker.blogspot.com/2009/10/telepresence-interoperability-part-3.html

PART 1: WHY TELEPRESENCE INTEROPERABILITY?

On October 6, 2009, Bob Dixon from OARnet moderated successful telepresence interoperability demonstration at the Fall Internet2 meeting in San Antonio, Texas. It included systems from Polycom, LifeSize, and Tandberg, and the short version of the story is in the joint press release http://finance.yahoo.com/news/Polycom-Internet2-OARnet-iw-1109370064.html?x=0&.v=1. While the memories from this event are still very fresh, I would like to spend some time and reflect on the long journey that led to this success.

First of all, why is telepresence interoperability so important?

The video industry is built on interoperability among systems from different vendors, and customers enjoy the ability to mix and match elements from Polycom, Tandberg, LifeSize, RadVision and other vendors in their video networks. As a result, video networks today rarely have equipment from only one vendor. It was therefore natural for the video community to strive for interoperability among multi-screen/multi-codec telepresence systems.

Most industry experts and visionaries in our industry subscribe to the idea that visual communication will become as pervasive as telephony today, and it has been widely recognized that the success of the good old Public Switch Telephone Network (PSTN) is based on vendors adhering to standards. Lack of interoperability, on the other hand, leads to inefficient network implementations of media gateways that transcode (translate) the digital audio and video information from one format to another thus increasing delay and decreasing quality. While gateways exist in voice networks, e.g. between PSTN and Voice over IP networks, their impact on delay and quality is far smaller than the impact of video gateways. Therefore, interoperability of video systems – telepresence and others – is even more important than interoperability of voice systems.

The International Multimedia Teleconferencing Consortium (IMTC) has traditionally driven interoperability based on the H.323 protocol. At the IMTC meeting in November’08 http://www.imtc.org/imwp/download.asp?ContentID=14027, the issue came up in three of the sessions and there were heated discussions how to tackle telepresence interoperability. The conclusion was that IMTC had expertise in signaling protocols (H.323) but not in the issues around multi-codec systems.

In February’09, fellow blogger John Bartlett wrote on NoJitter about the need for interoperability to enable business-to-business (B2B) telepresence and I replied on Video Networker http://videonetworker.blogspot.com/2009/03/business-to-business-telepresence.html, basically saying that proprietary mechanisms used in some telepresence systems create obstacles to interoperability.

In April’09, Bob Dixon from Ohio State and OARnet invited all telepresence vendors to the session ‘Telepresence Perspectives and Interoperability’ at the Spring Internet2 conference http://events.internet2.edu/2009/spring-mm/agenda.cfm?go=session&id=10000509&event=909. He chaired the session and, in conclusion, challenged all participating vendors to demonstrate interoperability of generally available products at the next Intrenet2 event. All vendors but HP were present. Initially, everyone agreed that this was a great idea. Using Internet2 to connect all systems would allow vendors to test without buying each others’ expensive telepresence systems. Bandwidth would not be an issue since Internet2 has so much of it. And since the interoperability would be driven by an independent third party, i.e. Bob Dixon, there would be no competitive fighting.

In June’09, I participated in the session ‘Interoperability: Separating Myth from Reality’ at the meeting of the Interactive Multimedia & Collaborative Communications Alliance (IMCCA) during InfoComm in Orlando, Florida http://www.infocommshow.org/infocomm2009/public/Content.aspx?ID=984&sortMenu=105005, and telepresence interoperability was on top of the agenda.

During InfoComm, Tandberg demonstrated connection between their T3 telepresence system and Polycom RPX telepresence system through the Tandberg Telepresence Server. The problem with such demos is always that you do not how much of it is real and how much is what we call ‘smoke and mirrors’. For those not familiar with this term, ‘smoke and mirrors’ refers to demos that are put together by modifying products and using extra wires, duct tape, glue and other high tech tools just to make it work for the duration of the demo. The main question I had around this demo was why a separate product like the Tandberg Telepresence Server was necessary? Couldn’t we just use a standard MCU with some additional layout control to achieve the same or even better results? To answer these questions, we needed an independent interoperability test. Ohio State, OARnet, and Internet2 would be the perfect vehicle for such test; they are independent and have a great reputation in the industry.

Stay tuned for Part 2 about the challenges to telepresence interoperability … http://videonetworker.blogspot.com/2009/10/part-2-telepresence-interoperability.html

Thursday, October 1, 2009

Cisco to Acquire Tandberg

Cisco announced today that they will acquire Tandberg, and this will have significant impact on the video communications market. It will reduce competition, and limit customers’ choices, especially in the telepresence space. It will, hurt Radvision who now fills the gap in Cisco’s video infrastructure portfolio.

I am however more concerned about the standards-compliance that have been the pillar of the video communication industry for years. Tandberg and Polycom worked together in international standardization bodies such as ITU-T and in industry consortiums such as IMTC to define standard mechanisms for video systems to communicate.

Cisco on the other hand is less interested in standards, and considers proprietary extensions as a way to gain competitive advantage. The concern of the video communication industry right now should be that the combined company will be so heavily dominated by Cisco that standards will become last priority, far after integrating Tandberg products with Cisco Call Manager and WebEx.

Telling is the fact that both Tandberg and Cisco declined participating in interoperability events over the last few months.

Wednesday, September 30, 2009

How to Manage Quality of Experience for Video?

Video calls require much higher network bandwidth than voice calls; they put therefore more strain on IP networks, and could overwhelm routers and switches to the point that they start losing packets. Video calls also tend to last longer than voice calls (the average length of a voice call is about 3 minutes); therefore, the probability that the network will experience performance degradation during a video call is higher. In addition to voice-related quality issues such as echo and noise, video struggles with freezes, artifacts, pixilation, etc.

So what can we do to guarantee high-quality user experience on video calls? This is an important question for organizations deploying on-premise video today. But due to the increased complexity of video networks, many organizations turn their video networks to managed service providers, and for them, measuring and controlling the quality of experience (QOE) is even more important. It allows managed SPs to identify and fix problems before the user calls the SP’ help desk; this impacts the SP’ bottom line directly.

Everyone who has used video long enough has encountered quality degradation at some point. Packet loss, jitter, and latency fluctuate depending on what else is being transmitted over the IP network. Quality of Service (QOS) mechanisms, such as DiffServ, help transmit real-time (video and voice) packets faster but even good QOS in the network does not necessarily mean that the user experience is good. QOE goes beyond just fixing network QOS; it also depends on the endpoints’ capability to compensate for network imperfections (through jitter buffers and packet recovery mechanisms), remove acoustic artifacts (like echo and noise), and combat image artifacts (like freezes and pixilation).

To monitor user experience, we can ask users to fill out a survey after every call. Skype, for example, is soliciting user feedback at the end of a call but how often do you fill out the form? And what if you are using a video endpoint with a remote control?

For longer video calls, it would be actually better if users report immediately when the issue happens, i.e. during the call. In practice, however, few users report problems while on a call. And even if they do, chances are that no one is available to investigate the issue immediately. In theory, the user could jot down the time when the problem happened and later ask the video network administrator to check if something happened in the IP network at that time. In reality, however, pressed by action items and back-to-back meetings, we just move on. As a result, problems do not get fixed and come back again and again.

Since we cannot rely on the users to report quality issues, we have to embed intelligence in the network itself to measure QOE and either make changes automatically to fix the problem (that would be the nirvana for every network manager) or at least create meaningful report identifying the problem area.

This technology exists today and has already been deployed in some Voice over IP networks. Most deployments use probes - small boxes distributed all over the network and inspecting RTP streams. Probes identify quality issues and report them to an aggregation tool that then generates reports for the network administrator. Integrating the probe’s functionality into endpoints makes the reports even more precise. For example, Polycom phones ship today with an embedded QOE agent that report to QOE management tools.

Originally developed for voice, QOE agents are getting more sophisticated, and now include some video capabilities. They can be used in video endpoints and multi-codec telepresence systems to monitor and report user experience. While this is currently not a priority for on-premise video deployments, QOE may become an important issue if more managed video services become available, as we all hope.

What can we expect to happen in this area in the future? The algorithms for calculating the impact of network issues on QOE will improve. Having the QOE agent embedded in the endpoint allows the endpoint’s application to submit additional quality information, e.g. echo parameters and noise level, to QOE management tools. This would give the tools more data points and lead to more precise identification of problems.

Since call success rate has direct impact on the quality of the user experience, one can expand the definition of QOE and use the same approach for monitoring call success rates. For example, the endpoint’s application can feed information about call success, failure, and reason for failure into the QOE agent and the agent can report that to the QOE reporting tool, which will detect lower-than-normal call success rates and alarm the network administrator. Call success rate can also be derived from Call Detail Records (CDRs) generated by the call control engine in the network; therefore, the alternative approach is to correlate the data from the CDRs with QOE reports from endpoints to identify issues.

While few organizations deploy QOE tools today we see increased interest among managed service providers who see value in any technology that allows them to avoid the dreaded help desk call. In the classic support scenario, a user complaint about bad call quality leads to finger pointing among voice SP, IP network SP, and organization’s IT department. Without proper tools, it is virtually impossible to identify the source of the problem. QOE reporting tools allow administrators to identify the source of the problem and are very valuable in distributed VOIP deployments.

In summation, QOE tools are new and still have a lot of room for improvement. However, the concept itself has been proven for voice and looks promising for video. In the future, look for wider support of QOE agents in voice and video products, and for wider deployment of QOE management tools, especially by managed service providers.

Monday, August 31, 2009

Is Flat Better than Flexible? The Curious Story of Resource Management in Conference Servers

When I joined the video communication industry in 2006, I learned that there is a huge argument in the industry about the best way to manage resources in a conference server (MCU). Three years later, the controversy continues and is a great topic for ‘Video Networker’.

Let’s start with the basics! Video endpoints connect to the conference server to join multi-point calls. The server has number of blades and each blade has a number of Digital Signaling Processors (DSPs) that process digital video. The ultimate flexibility for video users requires the server to transcode among video formats and to customize the Continuous Presence layout for each user. This flexibility costs a fair amount of resources which heavily depends on the quality of the processed video. Higher quality video means more information to process and requires more resources in the server. Not surprisingly, a conference servers can handle a smaller number of very high quality video connections (like HD 1080p), a larger number of high quality connections (like HD 720p), an even larger number of medium quality connections (like SD), and a huge number of low-quality video connections (like CIF). HD obviously stands for High Definition, SD - for Standard Definition, and CIF - for the lower quality Common Intermediate Format.

Having spent many years in the communications industry, this made perfect sense to me. Every server is more scalable when it has less work to do per user. In the case of a conference server, users connect at different quality depending on the capabilities of endpoints and the available network bandwidth. The conference server allocates resources to handle the new users dynamically, up until it runs out of resources and starts rejecting calls. In 2006, this was the way servers from Polycom, Tandberg and RadVision behaved, and there was not even a name for that behavior because it was natural.

Increased scalability was achieved in two ways. First, the video switching mode allowed server to avoid creating Continuous Presence screens. Only video from the loudest speaker was distributed to everybody else – very simple and scalable approach that led to reduced flexibility, and was totally inappropriate for many conferencing scenarios. A major limitation of video switching is that all sites must have the exact same capabilities (bit rate, resolution and frames per second), i.e., the conference server looks for a common denominator. One old video endpoint that can only support CIF resolution at 15fps takes the entire conference – including standard definition and high definition video endpoints - to CIF at 15fps. The second major drawback of video switching is that it only allows users to see ‘the loudest site’ on full screen. While it is nice to see the speaker on full screen, I feel very uncomfortable not seeing the body language of everybody else who is on the call. This limits the interactivity and negatively impacts the collaboration experience.

The second approach to scalability was ‘Conference on a Port’. The administrator of the conference server could select one Continuous Presence layout for the entire conference, and all participants who join received this layout. Again, the limited flexibility results in less work for the conference server (per user) and in increased scalability.

Back in 2006, I was in fact quite surprised to hear that a substantial number of people in the industry were excited by a new concept pushed by Codian and known as ‘a port is a port’ or ‘flat capacity’, which basically keeps the number of connections that the conference server supports constant, no matter whether the connection is HD, SD, or CIF. The proponents of this approach highlighted the simplicity of counting ports on servers. They also emphasized that, with the ‘flexible resource management’ approach, conference server administrators did not know for sure how many users the server can support. It is better, they said, to always have 20 ports rather than to have between 10 and 100 ports depending on connection types. Customers, they argued, should feel more comfortable buying a fixed number of ports.

So we had two competing philosophies in the market: ‘flexible resource management’ vs. ‘flat capacity’. The discussion went back and forth with urban legends coming from the ‘flat’ camp that new DSPs are somehow designed to perform better with ‘flat capacity’ and that there is so much performance on newer DSPs that you can afford to assign a lot of resources to a connection, no matter what quality it is. To the first argument, I know DSPs and they are designed to be a shared resource. Obviously, it is easier and simpler to assign a HD-capable DSP to a connection and let it process whatever quality comes in. It requires more sophisticated resource management to dynamically assign parts of DSPs to handle less demanding connections and full DSPs to HD connections. To the second argument, it is true that DSP performance increases but the complexity of handling HD is an order of magnitude higher than SD. Arguments for wasting resources sound hollow for conferencing servers that can cost $200,000 and up.

Anyway, the rational argumentation did not help resolve the discrepancies between ‘flat’ and ‘flexible’, and this resulted in new products that support both modes and allow the administrator to switch between them. For example, Polycom RMX 2000 easily switches between ‘flexible resource management’ and ‘fixed resource (‘flat capacity’) modes – the change does not even require restarting the server.

But now that the Pandora’s Box is opened, and everyone has an opinion on conference server resource management, there are a lot of new ideas for modes that make the server more efficient for certain applications. On the low end, desktop video is becoming popular and poses a new set of requirements to conference servers, so it is feasible to create a mode of operation dedicated to desktop video deployments. HD is less of an issue for desktop video but scalability is very important when entire organizations become video-enabled.

On the high end, multi-screen telepresence applications demand more performance per system from the conference server, while multiple video streams (one inbound and one outbound for each screen) must be associated and treated as a bundle. Some vendors like Tandberg decided to develop a completely separate product (Telepresence Server) to handle multi-point calls among multi-screen telepresence systems. I think this approach is an overreaction, some may say – an overkill. There are indeed some specific layouts that must be handled differently in a multi-screen telepresence environment but that does not mean putting a separate (and very expensive) server in the network just to handle telepresence calls. I think the approach where the standard conference server has a mode for multi-screen telepresence calls is much more sound from both business and technical perspectives – the main benefit is that you can still use the remaining resources on the server for regular calls among single-screen systems. This is also in accord with the maximum utilization philosophy driving ‘flexible resource management’.

As for the ‘separate telepresence server’ camp, it is not a coincidence that the same team that introduced ‘flat capacity’ is now pitching ‘separate telepresence server’. I see no innovation in limiting flexibility to achieve simplicity - true innovation is simplifying while keeping the flexibility intact.

In summation, conference servers are still the core of visual communication. In the past they had one application: video conferencing. Today, they have to handle video rooms, multi-screen telepresence systems, and desktop video. It is not surprising, therefore, that conference servers evolve and become more versatile. Adding new modes for resource management is a very pragmatic approach to satisfying requirements from new video applications, especially if switching among modes is fast and easy. Developing additional applications – which can run on general purpose computers and communicate with the conference server – is another valid approach. Developing separate servers for each application – room video, multi-screen telepresence, desktop video – is not scalable: it fills the network with hardware that is redundant in the bad way. As visual communication becomes mainstream and changes both our personal and professional lives, new sets of requirements to the conference server will emerge and the flexibility of the server platform to accommodate these requirements will decide whether conferencing servers will continue to be the heart of the video network or not.