The ATCA Summit http://www.advancedtcasummit.com/ (October 27-29, 2009) was a rare opportunity to think about and discuss the importance of hardware for our industry. As the communication industry becomes more software-driven, major industry events have focused on applications and solutions, and I rarely see good in-depth sessions on hardware. The ATCA Summit provided a refreshing new angle to communication technology.
First of all, ATCA stands for Advanced Telecom Computing Architecture and is a standard developed by the PICMG – a group of hardware vendors with great track record for defining solid hardware architectures: PCI, Compact PCI, MicroTCA, and ATCA. The ATCA Summit is the annual meeting of the ATCA community, or ecosystem, which includes vendors making chassis, blades, fans, power supplies, etc. components that can be used as tool kit to build a server quickly. Time-to-market is definitely an important reason companies turn to ATCA instead of developing their own hardware but equally important is that this telecom-grade (carrier-grade) hardware architecture provide very high scalability, redundancy, and reliability.
So, how does ATCA relate to visual communications? As visual communication becomes more pervasive and business critical both service providers (offering video services) and large enterprises (running their own video networks) start asking for more scalability and reliability in the video infrastructure. The core of the video infrastructure is the conference server (MCU), and the hardware architecture used in that network element has direct impact on the ability to support large video networks. HD video compression is very resource-intensive: raw HD video is about 1.5 gigabits per second, and modern H.264 compression technology can get it down to under 1 megabit per second. This 1500-fold compression requires powerful chips (DSPs) that generate a lot of heat; therefore, the conference server hardware must provide efficient power (think AC and DC power supplies) and cooling (think fans). But even in compressed form video is still using a lot of network bandwidth, and the conference server is the place where all video streams converge. Therefore, conference servers must have high input and output capabilities (think Gigabit Ethernet). Finally, some sort of blade architecture is required to allow for scalability, and server performance heavily depends on the way these blades are connected. The server component that connects the blades is referred to as ‘backplane’, although it does not need be physically in the back of the server. The ATCA architecture was built from ground up to meet these requirements. It was created with telecom applications in mind and has therefore high input/output, great power management and cooling, and a lot of mechanisms for high reliability.
The highlight of the ATCA Summit is always the Best of Show award. This year, Polycom RMX 4000 won Best of Show for infrastructure product, and I had the pleasure to receive the award. I posted a picture from the award ceremony here http://www.flickr.com/photos/20518315@N00/4080968072/. Subsequently, I presented in the session ‘The Users Talk Back’, and addressed the unique hardware functions in RMX 4000 that led to this award (http://www.flickr.com/photos/20518315@N00/4059010146/)
So, why did Polycom RMX 4000 win? I think it is mostly elegant engineering design and pragmatic decisions how to leverage standard hardware architecture to achieve unprecedented reliability. It starts with a high-throughput, low-overhead backplane (which we call ‘fabric switch’) that allows free flow of video across blades. This allows conferences to use resources from any of the blades. To illustrate the importance of this point, let’s briefly compare RMX 4000 to Tandberg MSE 8000 which combines 9 blades into a chassis but does not have a high-throughput backplane. Since video cannot flow freely among blades in MSE 8000, conferences are restricted to the resources available on a just one of the 9 blades. For example, if blade 1 supports 20 ports but 15 of them are already in use, you can only create a 5-party conference on that blade. If you need to start a 6-party conference, you cannot use blade 1, and have to look for another blade – let’s say blade 2 - that has 6 free ports. The 5 ports on blade 1 will stay idle until there is a conference of 5 or less participants. In fact, the ‘flat capacity’ software that is running on top of this hardware leads to even worse resource utilization on MSE 8000 but this article is about hardware, so I am not going into that subject (It is discussed in detail here http://videonetworker.blogspot.com/2009/08/curious-story-of-resource-management-in.html). The bottom line is that, with RMX 4000, you will be able to connect 5 participants to one blade and connect the sixth participant to another blade, without even noticing it.
Additional reliability can be gained by using DC power and full power supply redundancy. Direct Current (DC) power is used internally in all electronics equipment. However, power comes as Alternating Current (AC) over the power grid because AC power loss over long distances is lower than DC power loss. Once power reaches the data center, it makes sense to convert it once to DC and feed it to all servers, and that is why service providers and large enterprises running their own data centers like DC power. The alternative approach - provide AC power to each server and have each server convert it to DC - results in high conversion power loss, and is, basically, waste of energy, and should only be used if DC power is not available. RMX 4000 supports both AC and DC power but I am much more excited about the new DC power option. Each DC power supply has 1.5kW, and can power the entire RMX 4000. Best practice is to connect one DC power supply to the data center’s main power line and connect the second one to the battery array. Data centers have huge battery arrays that can keep them running even if the primary power line is down for hours or even days.
Reliability issues may arise from mixing media (video/audio) and signaling/management traffic, and therefore RMX 4000 completely separates these two types of traffic internally. This architectural approach also benefits security, since attacks against servers are usually about getting control of the signaling to manipulate the media. By clearly separating the two, RMX 4000 makes hijacking the server from outside impossible. Note that hijacking of voice conference servers is a major problem for voice service providers (I wrote about that here http://videonetworker.blogspot.com/2009/04/conferencing-service-providers-meet-at.html). As visual communication becomes more pervasive and business critical, similar issues can be expected in this space as well, and RMX 4000 is designed for that more dangerous future.
Finally, if a component in the conference server fails, it is critical that it can be replaced without disconnecting all calls and shutting down the server, thus preserving server-level reliability. All critical components in RMX 4000 are therefore hot swappable. This includes the four media blades (they are in the front of the chassis and host the video processing DSPs), RTM LAN modules (they are on the back of the chassis and connect to the IP network) and RTM ISDN modules (also on the back, connect to the ISDN network), power supplies, and fans. Each of these components can be removed and replaced with a new one while the RMX 4000 server is running.
..
I will discuss the topics of network-based redundancy and reliability in a separate article. Stay tuned!
This blog discusses collaboration market and technologies including video conferencing, web conferencing, and team collaboration tools.
Showing posts with label conference server. Show all posts
Showing posts with label conference server. Show all posts
Monday, November 9, 2009
Monday, August 31, 2009
Is Flat Better than Flexible? The Curious Story of Resource Management in Conference Servers
When I joined the video communication industry in 2006, I learned that there is a huge argument in the industry about the best way to manage resources in a conference server (MCU). Three years later, the controversy continues and is a great topic for ‘Video Networker’.
Let’s start with the basics! Video endpoints connect to the conference server to join multi-point calls. The server has number of blades and each blade has a number of Digital Signaling Processors (DSPs) that process digital video. The ultimate flexibility for video users requires the server to transcode among video formats and to customize the Continuous Presence layout for each user. This flexibility costs a fair amount of resources which heavily depends on the quality of the processed video. Higher quality video means more information to process and requires more resources in the server. Not surprisingly, a conference servers can handle a smaller number of very high quality video connections (like HD 1080p), a larger number of high quality connections (like HD 720p), an even larger number of medium quality connections (like SD), and a huge number of low-quality video connections (like CIF). HD obviously stands for High Definition, SD - for Standard Definition, and CIF - for the lower quality Common Intermediate Format.
Having spent many years in the communications industry, this made perfect sense to me. Every server is more scalable when it has less work to do per user. In the case of a conference server, users connect at different quality depending on the capabilities of endpoints and the available network bandwidth. The conference server allocates resources to handle the new users dynamically, up until it runs out of resources and starts rejecting calls. In 2006, this was the way servers from Polycom, Tandberg and RadVision behaved, and there was not even a name for that behavior because it was natural.
Increased scalability was achieved in two ways. First, the video switching mode allowed server to avoid creating Continuous Presence screens. Only video from the loudest speaker was distributed to everybody else – very simple and scalable approach that led to reduced flexibility, and was totally inappropriate for many conferencing scenarios. A major limitation of video switching is that all sites must have the exact same capabilities (bit rate, resolution and frames per second), i.e., the conference server looks for a common denominator. One old video endpoint that can only support CIF resolution at 15fps takes the entire conference – including standard definition and high definition video endpoints - to CIF at 15fps. The second major drawback of video switching is that it only allows users to see ‘the loudest site’ on full screen. While it is nice to see the speaker on full screen, I feel very uncomfortable not seeing the body language of everybody else who is on the call. This limits the interactivity and negatively impacts the collaboration experience.
The second approach to scalability was ‘Conference on a Port’. The administrator of the conference server could select one Continuous Presence layout for the entire conference, and all participants who join received this layout. Again, the limited flexibility results in less work for the conference server (per user) and in increased scalability.
Back in 2006, I was in fact quite surprised to hear that a substantial number of people in the industry were excited by a new concept pushed by Codian and known as ‘a port is a port’ or ‘flat capacity’, which basically keeps the number of connections that the conference server supports constant, no matter whether the connection is HD, SD, or CIF. The proponents of this approach highlighted the simplicity of counting ports on servers. They also emphasized that, with the ‘flexible resource management’ approach, conference server administrators did not know for sure how many users the server can support. It is better, they said, to always have 20 ports rather than to have between 10 and 100 ports depending on connection types. Customers, they argued, should feel more comfortable buying a fixed number of ports.
So we had two competing philosophies in the market: ‘flexible resource management’ vs. ‘flat capacity’. The discussion went back and forth with urban legends coming from the ‘flat’ camp that new DSPs are somehow designed to perform better with ‘flat capacity’ and that there is so much performance on newer DSPs that you can afford to assign a lot of resources to a connection, no matter what quality it is. To the first argument, I know DSPs and they are designed to be a shared resource. Obviously, it is easier and simpler to assign a HD-capable DSP to a connection and let it process whatever quality comes in. It requires more sophisticated resource management to dynamically assign parts of DSPs to handle less demanding connections and full DSPs to HD connections. To the second argument, it is true that DSP performance increases but the complexity of handling HD is an order of magnitude higher than SD. Arguments for wasting resources sound hollow for conferencing servers that can cost $200,000 and up.
Anyway, the rational argumentation did not help resolve the discrepancies between ‘flat’ and ‘flexible’, and this resulted in new products that support both modes and allow the administrator to switch between them. For example, Polycom RMX 2000 easily switches between ‘flexible resource management’ and ‘fixed resource (‘flat capacity’) modes – the change does not even require restarting the server.
But now that the Pandora’s Box is opened, and everyone has an opinion on conference server resource management, there are a lot of new ideas for modes that make the server more efficient for certain applications. On the low end, desktop video is becoming popular and poses a new set of requirements to conference servers, so it is feasible to create a mode of operation dedicated to desktop video deployments. HD is less of an issue for desktop video but scalability is very important when entire organizations become video-enabled.
On the high end, multi-screen telepresence applications demand more performance per system from the conference server, while multiple video streams (one inbound and one outbound for each screen) must be associated and treated as a bundle. Some vendors like Tandberg decided to develop a completely separate product (Telepresence Server) to handle multi-point calls among multi-screen telepresence systems. I think this approach is an overreaction, some may say – an overkill. There are indeed some specific layouts that must be handled differently in a multi-screen telepresence environment but that does not mean putting a separate (and very expensive) server in the network just to handle telepresence calls. I think the approach where the standard conference server has a mode for multi-screen telepresence calls is much more sound from both business and technical perspectives – the main benefit is that you can still use the remaining resources on the server for regular calls among single-screen systems. This is also in accord with the maximum utilization philosophy driving ‘flexible resource management’.
As for the ‘separate telepresence server’ camp, it is not a coincidence that the same team that introduced ‘flat capacity’ is now pitching ‘separate telepresence server’. I see no innovation in limiting flexibility to achieve simplicity - true innovation is simplifying while keeping the flexibility intact.
In summation, conference servers are still the core of visual communication. In the past they had one application: video conferencing. Today, they have to handle video rooms, multi-screen telepresence systems, and desktop video. It is not surprising, therefore, that conference servers evolve and become more versatile. Adding new modes for resource management is a very pragmatic approach to satisfying requirements from new video applications, especially if switching among modes is fast and easy. Developing additional applications – which can run on general purpose computers and communicate with the conference server – is another valid approach. Developing separate servers for each application – room video, multi-screen telepresence, desktop video – is not scalable: it fills the network with hardware that is redundant in the bad way. As visual communication becomes mainstream and changes both our personal and professional lives, new sets of requirements to the conference server will emerge and the flexibility of the server platform to accommodate these requirements will decide whether conferencing servers will continue to be the heart of the video network or not.
Let’s start with the basics! Video endpoints connect to the conference server to join multi-point calls. The server has number of blades and each blade has a number of Digital Signaling Processors (DSPs) that process digital video. The ultimate flexibility for video users requires the server to transcode among video formats and to customize the Continuous Presence layout for each user. This flexibility costs a fair amount of resources which heavily depends on the quality of the processed video. Higher quality video means more information to process and requires more resources in the server. Not surprisingly, a conference servers can handle a smaller number of very high quality video connections (like HD 1080p), a larger number of high quality connections (like HD 720p), an even larger number of medium quality connections (like SD), and a huge number of low-quality video connections (like CIF). HD obviously stands for High Definition, SD - for Standard Definition, and CIF - for the lower quality Common Intermediate Format.
Having spent many years in the communications industry, this made perfect sense to me. Every server is more scalable when it has less work to do per user. In the case of a conference server, users connect at different quality depending on the capabilities of endpoints and the available network bandwidth. The conference server allocates resources to handle the new users dynamically, up until it runs out of resources and starts rejecting calls. In 2006, this was the way servers from Polycom, Tandberg and RadVision behaved, and there was not even a name for that behavior because it was natural.
Increased scalability was achieved in two ways. First, the video switching mode allowed server to avoid creating Continuous Presence screens. Only video from the loudest speaker was distributed to everybody else – very simple and scalable approach that led to reduced flexibility, and was totally inappropriate for many conferencing scenarios. A major limitation of video switching is that all sites must have the exact same capabilities (bit rate, resolution and frames per second), i.e., the conference server looks for a common denominator. One old video endpoint that can only support CIF resolution at 15fps takes the entire conference – including standard definition and high definition video endpoints - to CIF at 15fps. The second major drawback of video switching is that it only allows users to see ‘the loudest site’ on full screen. While it is nice to see the speaker on full screen, I feel very uncomfortable not seeing the body language of everybody else who is on the call. This limits the interactivity and negatively impacts the collaboration experience.
The second approach to scalability was ‘Conference on a Port’. The administrator of the conference server could select one Continuous Presence layout for the entire conference, and all participants who join received this layout. Again, the limited flexibility results in less work for the conference server (per user) and in increased scalability.
Back in 2006, I was in fact quite surprised to hear that a substantial number of people in the industry were excited by a new concept pushed by Codian and known as ‘a port is a port’ or ‘flat capacity’, which basically keeps the number of connections that the conference server supports constant, no matter whether the connection is HD, SD, or CIF. The proponents of this approach highlighted the simplicity of counting ports on servers. They also emphasized that, with the ‘flexible resource management’ approach, conference server administrators did not know for sure how many users the server can support. It is better, they said, to always have 20 ports rather than to have between 10 and 100 ports depending on connection types. Customers, they argued, should feel more comfortable buying a fixed number of ports.
So we had two competing philosophies in the market: ‘flexible resource management’ vs. ‘flat capacity’. The discussion went back and forth with urban legends coming from the ‘flat’ camp that new DSPs are somehow designed to perform better with ‘flat capacity’ and that there is so much performance on newer DSPs that you can afford to assign a lot of resources to a connection, no matter what quality it is. To the first argument, I know DSPs and they are designed to be a shared resource. Obviously, it is easier and simpler to assign a HD-capable DSP to a connection and let it process whatever quality comes in. It requires more sophisticated resource management to dynamically assign parts of DSPs to handle less demanding connections and full DSPs to HD connections. To the second argument, it is true that DSP performance increases but the complexity of handling HD is an order of magnitude higher than SD. Arguments for wasting resources sound hollow for conferencing servers that can cost $200,000 and up.
Anyway, the rational argumentation did not help resolve the discrepancies between ‘flat’ and ‘flexible’, and this resulted in new products that support both modes and allow the administrator to switch between them. For example, Polycom RMX 2000 easily switches between ‘flexible resource management’ and ‘fixed resource (‘flat capacity’) modes – the change does not even require restarting the server.
But now that the Pandora’s Box is opened, and everyone has an opinion on conference server resource management, there are a lot of new ideas for modes that make the server more efficient for certain applications. On the low end, desktop video is becoming popular and poses a new set of requirements to conference servers, so it is feasible to create a mode of operation dedicated to desktop video deployments. HD is less of an issue for desktop video but scalability is very important when entire organizations become video-enabled.
On the high end, multi-screen telepresence applications demand more performance per system from the conference server, while multiple video streams (one inbound and one outbound for each screen) must be associated and treated as a bundle. Some vendors like Tandberg decided to develop a completely separate product (Telepresence Server) to handle multi-point calls among multi-screen telepresence systems. I think this approach is an overreaction, some may say – an overkill. There are indeed some specific layouts that must be handled differently in a multi-screen telepresence environment but that does not mean putting a separate (and very expensive) server in the network just to handle telepresence calls. I think the approach where the standard conference server has a mode for multi-screen telepresence calls is much more sound from both business and technical perspectives – the main benefit is that you can still use the remaining resources on the server for regular calls among single-screen systems. This is also in accord with the maximum utilization philosophy driving ‘flexible resource management’.
As for the ‘separate telepresence server’ camp, it is not a coincidence that the same team that introduced ‘flat capacity’ is now pitching ‘separate telepresence server’. I see no innovation in limiting flexibility to achieve simplicity - true innovation is simplifying while keeping the flexibility intact.
In summation, conference servers are still the core of visual communication. In the past they had one application: video conferencing. Today, they have to handle video rooms, multi-screen telepresence systems, and desktop video. It is not surprising, therefore, that conference servers evolve and become more versatile. Adding new modes for resource management is a very pragmatic approach to satisfying requirements from new video applications, especially if switching among modes is fast and easy. Developing additional applications – which can run on general purpose computers and communicate with the conference server – is another valid approach. Developing separate servers for each application – room video, multi-screen telepresence, desktop video – is not scalable: it fills the network with hardware that is redundant in the bad way. As visual communication becomes mainstream and changes both our personal and professional lives, new sets of requirements to the conference server will emerge and the flexibility of the server platform to accommodate these requirements will decide whether conferencing servers will continue to be the heart of the video network or not.
Thursday, March 5, 2009
Scalable video conference servers
A lot of the discussion in the video industry these days is around video conference servers (also known as bridges and MCUs). With the advances of video communication technology and the deployment of HD video the load on video conference servers is growing because they have to process more bits per second to support HD video calls. In addition, desktop video deployments rapidly increase the size of video networks.
Fundamentally, there are three ways to make video conference servers more scalable. The first one is to build a large server, use carrier-grade architecture and try to squeeze in as much computing power as possible into a large chassis. This is the approach Tandberg is taking with the MSE 8000. The benefit of such approach is that it easy to explain to resellers, integrators and customers: ‘you had a small server, now you are running out of resources, so get a big one’. The disadvantage is that the server becomes an extremely critical single point of failure; if the server is down, or its part of the IP network is down, the entire video service is impacted. There is also the cost aspect - buying such large server is a considerable chunk of money – but I am looking at it from a network design perspective and can only say that it is impossible to find an optimum location for such server in the network. Enterprise, government, education and health networks are all so distributed these days that placing the server in any one location leads to inefficient use of the network bandwidth and decreased quality for participants from other locations.
The second approach to scalability is to build a conference server sufficient for mid-sized video deployments and create a new architecture that allows you to combine many such conference servers into one pool of conferencing resources – to meet the needs of large organizations. You can increase this pool by adding conference servers and decrease it by removing them. The management server that manages all resources reroutes video calls to the most appropriate resource in the entire network. You can make the selection algorithm as sophisticated as you want, e.g. the algorithm may select the conference server that is closest to the majority of participants, or select the server that has the horsepower to support the quality that the participants require for that particular call. The benefit of this architecture is that you can spread the conference servers across your networks – thus avoiding bottlenecks and congestions - and still manage all servers as one giant virtual conference server. This is Polycom’s architecture: the conference server is RMX 2000; the resource management server is DMA 7000. Networking experts understand the high reliability and survivability of this approach – distributed computing and load balancing have been the preferred way to achieve scalability of applications for long time. The real challenge with this approach is to educate traditional video equipment resellers and integrators who look at the conference server as ‘the bridge’ (i.e. one box) and not as a service that can and must be distributed across the network to scale.
The third approach is to completely change the distribution of computing power between endpoints and conference servers, i.e. move more of the computing to the video endpoints (requires more powerful/expensive hardware for endpoints) and reduce the computing power in the conference server. This is the approach which the startup Vidyo is taking. Simplifying the conference server is a great idea but it remains to be seen if this benefit can outweigh the need for more performance in the endpoints. More importantly, this approach is incompatible with the installed base of video equipment and requires signaling and media gateways for interoperability. Gateways – and especially media gateways – introduce delays, decrease video and audio quality, and add substantial cost to the solution.
You can find more details about the scalability mechanisms discussed in this posting, as well as diagrams explaining the configurations, in the white paper ‘Scalable Infrastructure for Distributed Video’ (see link in the ‘White Papers and Articles’ section below on this blog). I will also address this subject in my presentation ‘Visual Communication – Believe the hype and prepare for the impact’ at InfoComm (see link in the ‘Speaking Engagements’ section below).
Fundamentally, there are three ways to make video conference servers more scalable. The first one is to build a large server, use carrier-grade architecture and try to squeeze in as much computing power as possible into a large chassis. This is the approach Tandberg is taking with the MSE 8000. The benefit of such approach is that it easy to explain to resellers, integrators and customers: ‘you had a small server, now you are running out of resources, so get a big one’. The disadvantage is that the server becomes an extremely critical single point of failure; if the server is down, or its part of the IP network is down, the entire video service is impacted. There is also the cost aspect - buying such large server is a considerable chunk of money – but I am looking at it from a network design perspective and can only say that it is impossible to find an optimum location for such server in the network. Enterprise, government, education and health networks are all so distributed these days that placing the server in any one location leads to inefficient use of the network bandwidth and decreased quality for participants from other locations.
The second approach to scalability is to build a conference server sufficient for mid-sized video deployments and create a new architecture that allows you to combine many such conference servers into one pool of conferencing resources – to meet the needs of large organizations. You can increase this pool by adding conference servers and decrease it by removing them. The management server that manages all resources reroutes video calls to the most appropriate resource in the entire network. You can make the selection algorithm as sophisticated as you want, e.g. the algorithm may select the conference server that is closest to the majority of participants, or select the server that has the horsepower to support the quality that the participants require for that particular call. The benefit of this architecture is that you can spread the conference servers across your networks – thus avoiding bottlenecks and congestions - and still manage all servers as one giant virtual conference server. This is Polycom’s architecture: the conference server is RMX 2000; the resource management server is DMA 7000. Networking experts understand the high reliability and survivability of this approach – distributed computing and load balancing have been the preferred way to achieve scalability of applications for long time. The real challenge with this approach is to educate traditional video equipment resellers and integrators who look at the conference server as ‘the bridge’ (i.e. one box) and not as a service that can and must be distributed across the network to scale.
The third approach is to completely change the distribution of computing power between endpoints and conference servers, i.e. move more of the computing to the video endpoints (requires more powerful/expensive hardware for endpoints) and reduce the computing power in the conference server. This is the approach which the startup Vidyo is taking. Simplifying the conference server is a great idea but it remains to be seen if this benefit can outweigh the need for more performance in the endpoints. More importantly, this approach is incompatible with the installed base of video equipment and requires signaling and media gateways for interoperability. Gateways – and especially media gateways – introduce delays, decrease video and audio quality, and add substantial cost to the solution.
You can find more details about the scalability mechanisms discussed in this posting, as well as diagrams explaining the configurations, in the white paper ‘Scalable Infrastructure for Distributed Video’ (see link in the ‘White Papers and Articles’ section below on this blog). I will also address this subject in my presentation ‘Visual Communication – Believe the hype and prepare for the impact’ at InfoComm (see link in the ‘Speaking Engagements’ section below).
Subscribe to:
Posts (Atom)