|
June 1997
By Amy K. Larsen
Monitoring Tools
Guaranteed Service?
Reporting tools and service packages can help in setting up SLAs, but it's still not an easy task
|
|
Forgive sna vets for waxing nostalgic about the good old days. Among the big benefits afforded by big iron was the ability to set service-level agreements with various departments. Verify usage. Charge
for bandwidth. Prove performance. It was all so easy ... because it was centralized. Setting SLAs on today's distributed nets makes for a whole new set of challenges. But that's where today's device and application reporting tools could help. Nearly 20 vendors are shipping products that they say can be used to establish SLAs with internal departments and outside service providers. Originally built for troubleshooting and capacity planning, these tools are taking on a new role as net managers seek to verify performance across the enterprise. "We aren't a profit center," says Alan Robson, director of network systems for Cox Target Media Inc. (Largo, Fla.), a direct mail order business. "So we have to prove we're doing a good job."
But before writing up those SLAs, it pays to learn the ABCs. There
are a lot of products on the market, and every one of them can report
on one or more of three basic service-level metrics: availability,
reliability, and response time. So networkers will have to look deeper
in
to the differentiators. All but two packages compile performance
trends based on RMON (remote monitoring) and SNMP stats, many support
vendor extensions, and a handful use proprietary agent technology.
While most products can monitor WAN links, only a few focus on the
LAN. Scalability also sets products apart: Some can monitor as many as
5,000 devices, while others are considerably more limited.
Report-generation capabilities vary, too--a key concern given that
SLAs are often used for service provisioning and billing purposes.
Price also can prove a big differentiator.
And even then, there's more to SLAs than technology. "There are
lots of techie metrics we can apply," says Sydney Fisher, a senior
product manager for Network General Corp. (Menlo Park, Calif.). "But a
company still has to set its own business goals and pick relevant
performance stats to measure its service levels against." In other
words, it means customization--something that a company short on
skilled staff might not want (or be able
) to deal with. Also, these
are performance-monitoring applications only--they don't deliver fault
information.
For some companies, farming out the service-level monitoring chores
may be the answer instead. Right now, three providers offer such
services, and they rely on the same standards-based collecting methods
most products use. While ongoing fees can make these services pricier
than off-the-shelf packages, the money saved by not having to staff up
can usually offset the cost.
Agreement Basics
When it comes to evaluating service levels, net managers can use
three broad parameters: availability, reliability, and response time.
But it's important to understand the differences among them.
|
 BUILD YOUR OWN CUSTOM TABLE
Global Coverage
|
The bulk of vendors--including Cisco Systems Inc. (San Jose,
Calif.), Kaspia Systems Inc. (Beaverton, Ore.), Network General, and
Visual Networks Inc. (Rockville, Md.)--use availability, reliability,
or a combination (see
Table 1
). There is, however, something to keep
in mind. Availability refers to uptime (of either a device or the
network itself), while reliability refers to how often a device or
network goes down or how long it stays down--in other words, the mean
time between failure and repair. And generally, reliability makes for
a better metric since it measures consistency. Say a router port is
out of commission 1 minute out of every 10. Availability is 90
percent, but in fact the port is consistently unreliable.
The other metric is response time. While availability and
reliability offer high-level views of performance, response time is a
way to gauge how the end-user is affected. There are nine vendors
whose produ
cts use this parameter, including Platinum Technology Inc.
(Oakbrook Terrace, Ill.), Solcom Systems Inc. (Reston, Va.), and
Technically Elite (San Jose, Calif.).
Net managers also have to figure out how (or whether) these
parameters will fit into the SLA terms they set. That means taking
care of two crucial, non-technical details: defining performance
objectives and the price of meeting them. In other words, terms need
to match goals.
"What we need is a metric that ties our technology model to our
business model," says Cox Target's Robson. He's looking at the SLA
Conformance Manager from Infovista Corp. (Redwood City, Calif.)--which
uses all three parameters--to help his IT group set SLAs with other
divisions of the direct mail company. "We need something to tell us
what kind of application response times we're getting on our
order-entry application. We have to show management how at the
infrastructure level we're helping get more coupons out the door."
Goldman Sachs Group L.P. (New Yo
rk) had a different
objective--tracking the service level of the IT staff itself. The
problem? Highly skilled database administrators, system engineers, and
developers were responding to problems operations personnel were paid
to manage--and timely troubleshooting of normal network problems came
into question. "We wanted to prove we could handle the load," says Hal
Uygur, vice president of technical operations/enterprise. "So we set
up an SLA guaranteeing we would respond to a problem within two
minutes of receiving an alert."
The investment firm then used a combination of customized and
off-the-shelf monitoring products, management systems, and event
correlation software to track how long it took for operators to
respond to alerts. The results showed that the department met its
two-minute goal 85 percent of the time. Goldman Sachs also devised a
rating system to determine how operators solved problems and whether
they needed assistance. Ultimately, Uygur wants to use this
information to identify spec
ific skills--which means that specific
problems can be directed to the person most capable of handling
them.
The upshot? Network managers shouldn't simply rely on technology
when it comes to setting up SLAs. They must also figure out how the
tools can be used to meet their specific objectives.
Call My Agent
Of course, the ability to prove performance means tracking just
what's happening on the network. Monitoring tools use either an RMON
(remote monitoring) probe, an SNMP agent, or a proprietary data
collector to gather network and application performance stats. All
products rely on a software-based polling agent, but it's the data
source that determines what kind of report the monitoring application
will generate.
For instance, almost every package uses traditional console
architecture to poll SNMP-managed devices and generate performance
trend reports. And of those, most support SNMP MIBs (management
information bases) I and II, which means they can gather hub, switch,
and
server data--as well as router interface stats. This information
is then aggregated to quantify such network conditions as device
availability and utilization.
Fewer packages gather data from RMON probes. But that's probably
because RMON isn't as common a feature in corporate nets as SNMP, so
vendors don't build in such capabilities. For those net managers who
have deployed RMON, however, it might then be worthwhile to check just
how much RMON support is available. For instance, is a tool limited to
gathering network layer stats (RMON I)? Or can it track conversations
from end-station to host (RMON II)?
There's a practical reason for finding out: The more extensive the
data, the more thorough the SLA. "The detail in our reports really
depends on how richly instrumented the network is," says Network
General's Fisher. "If there are only RMON I probes, we can list top
talkers, but we can't identify top conversations." Without that
information, net managers can't accurately bill departments for
ba
ndwidth usage.
This doesn't mean that net managers have to rush out and install
RMON probes end to end (after all, a four-port Ethernet probe--which
typically costs $7,000--can make a big dent in the budget). Putting
probes in the right places could suffice. "If the probes are installed
at strategic points on the network, you don't need a probe on every
segment," says Jim McQuaid, director of business development for
Netscout Systems Inc. (Chelmsford, Mass.). The vendor recommends
installing RMON II agents in a centralized server farm close to a WAN
link, so that traffic volume and inbound and outbound conversations
can be monitored.
And with hardware vendors now selling RMON-equipped hubs and
switches, there's even less of a need to purchase standalone probes
for every critical segment. Products that can cull information from
RMON agents in network gear include Bestview from BGS Systems Inc.
(Waltham, Mass.), Network Health from Concord Communications Inc.
(Marlborough, Mass.), and TrendSNMP
from Desktalk Systems Inc.
(Torrance, Calif.).
Request an Extension
Still, no matter how "well-instrumented" a network, the performance
statistics yielded by SNMP MIBs and RMON probes don't come close to
matching the level of service data furnished by the operating system
of a mainframe. "The tool to put together mainframe-like SLAs for
distributed networks would have to use something akin to artificial
intelligence," says Tim Riley, director of Optivity product marketing
for Bay Networks Inc. (Santa Clara, Calif.). "It's difficult to
productize that."
That's why most analysis tool vendors include proprietary
extensions with their products. It's just something that gives
networkers another layer of detail. For instance, with SNMP router
extensions for Bay and Cisco, Concord's Network Health can furnish
such information as router CPU utilization, which can then be used as
part of the performance measurement.
Other vendors have similar plans. Kaspia, for example, has
approached
firewall vendors about using data from log files for
performance information. Netscout and Technically Elite have added
proprietary extensions to record SQL requests and responses so they
can measure database application response times.
In all, 17 applications and service packages use private vendor
extensions or proprietary collection agents to gather performance
statistics. Cisco's Netsys Tools, Network General's Service Level
Manager, and Visual Networks' Visual Onramp and Visual Uptime all use
protocol decodes from an analyzer to build performance reports. The
Cisco and Network General products both incorporate data gathered by
Network General's Sniffer systems, while Visual Networks uses a
combination of SNMP MIB data and a protocol analyzer/probe built into
its CSU/DSU to decode packet information from the WAN link.
|
 MESSAGE SERVER
Join the discussion about SLA's
|
Application performance monitoring tools--like Ecoscope from
Compuware Corp. (Farmington Hills, Mich.), HP's Netmetrix Reporter,
System Management Architecture (SMA) from Jyra Research Inc. (San
Jose, Calif.), Application Expert from Optimal Networks Corp. (Palo
Alto, Calif.), and Platinum's Wiretap--rely in part on nonstandard
agents that passively monitor network activity and supply
response-time information. Optimal's Application Expert also can
follow application threads and record response times between servers,
as well as the total round-trip response time.
Jyra's SMA uses two methods to measure response time: traceroute
and synthetic transaction monitoring. Traceroute, a mechanism of ICMP
(Internet control message protocol), tests performance by sending a
traffic stream and recording the application response time. Synthetic
transaction mo
nitoring adds a request for time-stamped transaction
information to an actual application traffic stream leaving an
end-station. Once the transaction is finished, the application is
notified of total round-trip response time. Jyra says that with this
capability, its product can spot the source of delays even on a
carrier network (see "Added Insight Into Carrier Networks," April
1997).
HP's Netmetrix Reporter uses information from several sources to
compile application response times. Besides taking in data from RMON
probes, it uses Measureware agents to collect database request and
reply times and operating system statistics. Netmetrix Reporter also
can read information from applications built using the ARM
(Application Response Measurement) application program interface.
ARM--which is the result of a product initiative between IBM/Tivoli
Systems Inc. (Austin, Texas) and HP--permits an application to be
time-stamped as it passes through different devices and hosts. (ARM
has not been widely implemented
by third-party software vendors, but
application developers can download it from the Tivoli Web site:
http://www.tivoli.com /tivevery/download.html.)
The Wide Side
Net managers also have to decide if they want to limit SLAs just to
the LAN--or whether they want to use them to keep tabs on service
providers as well.
More than a dozen products and services can be used to monitor WAN
link activity. They deliver data that can be used to verify contracts
between a corporation and its carrier, be it frame relay, leased line,
or ISP (Internet service provider). This information also can be used
to validate internal SLAs that are based in part on WAN
performance.
Establishing an SLA with a service provider is a tricky business.
After all, carriers aren't exactly eager to enter into them. "SLAs
catch carriers with their proverbial pants down," says Jim Parkhurst,
a senior staff engineer at MCI Communications Corp. (Washington,
D.C.).
Net managers also have to remember that so
me problems fall outside
the carrier's purview. Getting access to the ISP network, for example,
means going through the local loop--a no-man's land in terms of
automatic visibility. Although carriers sometimes furnish
service-level reports to their customers, net managers would be better
served by monitoring performance themselves. They can use specific
monitoring products to do this, or they can choose from the network
monitoring services offered by Centron DPL Co. (Eden Prairie, Minn.),
International Network Services Inc. (INS, Sunnyvale, Calif.), or
Netops Corp (New Fairfield, Conn.).
Either way, the best place to start is by matching the service
level to a business application's technical requirements, says MCI's
Parkhurst. For instance, setting a CIR (committed information rate)
for frame relay service depends on more than basic usage
requirements.
Net managers also need to know just what kind of delays various
applications can live with. Isochronous protocols, for example, can
handle onl
y brief delays before timeouts lead to session loss. Thus
net managers have to verify that delays are lower than 150
milliseconds when sending SNA traffic over frame relay.
Packages have different ways of monitoring WAN service levels.
Netclarity from Ascend Communications Inc. (Alameda, Calif.), Cisco's
Netsys Tools, Concord's Network Health, Desktalk's TrendSNMP,
Infovista's SLA Conformance Management, Network General's Routerpm,
and Visual Networks' Visual Uptime all measure delay between frames on
each PVC (permanent virtual circuit) and then compile a report. Visual
Networks uses a proprietary probe to report on frame delay.
Using a product like Ascend's Netclarity or Network General's
Routerpm, it's also possible to verify availability of a certain
circuit. Net managers can poll the attached router's MIB tables at
regular intervals to see if the circuit is transmitting, idle, or
down.
Some tools can also be used to monitor leased-line performance.
Among these are Ascend's Netclar
ity, Cisco's Netsys Tools, Concord's
Network Health, Infovista's SLA Conformance Management and Visual
Networks' Visual Uptime, which measure serial-link efficiency, or the
throughput of a serial link in relation to the number of dropped
packets.
Numbers Game
There's something else net managers should keep in mind when
purchasing a reporting package: the number of devices it can monitor.
SNMP's scaling limitations and net managers' storage constraints can
conspire to limit monitoring capabilities to as few as 200 devices.
But other vendors say their products can handle as many as 5,000.
(Actually, they point out that there is theoretically no limit to the
number of devices they cover; the numbers given refer to the largest
real-world installation a single copy of the tool is known to
monitor.)
Data export and import capability is something else net managers
should explore. It can come in particularly handy when using a
correlation system (or "manager of managers")--such as Netcool/Om
nibus
from Micromuse Inc. (San Francisco)--to tie together data from
different management and monitoring systems. Twelve products,
including Ascend's Netclarity, BGS Systems' Bestview, and Compuware's
Ecoscope, offer some type of import/export capability to third-party
applications.
Dollar Signs
Of course, formal SLAs define specific financial consequences for
service that's not up to snuff. "These individuals are not
troubleshooting networks on a daily basis," says Rob Markovitz, vice
president of marketing for Visual Networks. "But they need to make
sure they are getting the $20 million worth of service they've paid
for."
Enforcing such agreements means coming up with specific proof,
either through an executive-level summary or a usage-based cost
analysis. Unfortunately, only seven products and services issue the
kind of non-technical reports that chief financial officers can use to
hold carriers' feet to the fire or bill departments for over-use. The
rest say they'll offer such ca
pabilities by the end of the year.
Of the products that offer reports, those from Ascend, Concord,
Infovista, INS, and Visual Networks also help in bandwidth planning.
Concord's, Netscout's, and Visual Networks' also offer usage-based
bandwidth accounts.
It's worth noting that these executive report capabilities can come
at a premium. Both Concord and Visual Networks are charging an extra
$10,000 for financial report summaries. But other vendors, Ascend and
Netscout among them, include these reports as part of the basic
package.
|
 CONTACT AUTHOR
alarsen@data.com
|
Speaking of which: Base package prices can vary from as little as
$900 to as much as $80,000. Of course, with a range that wide, it's
pretty obvious that net managers will get what they pay for. The $900
tool, Clear Stat
s Lite from Clear Systems Inc. (Irving, Texas), covers
only 200 LAN devices and furnishes only a limited number of
rudimentary reports on device availability.
But high-end packages like Desktalk's TrendSNMP, Concord's Network
Health, Kaspia's Network Monitoring System, and Network General's
Routerpm, on the other hand, can track service levels on frame relay.
Cisco's Netsys Tools goes beyond basic performance reporting to
suggest corrective measures for solving router configuration
problems.
And there's one final thing to remember: Prices for shrink-wrapped
products don't include hardware costs or the cost of hiring personnel
to field key stats from the applications.
Amy K. Larsen is LANs/ network management editor for Data Communications. She can be reached at
alarsen@data.com
.
[
Home
]
[
Registration
|
Subscriptions
]
[
Contact Us
|
E-Mail
]
|
|
|
 |
 |
|