| ★ wanayoo — archive 1999 http://www.samizdat.com/script/av14.htm | Nouvelle recherche | Portail wanayoo |
Keep in mind that AltaVista doesn't index everything.
Actually, the way this works is much to the advantage of small companies and individuals. It wasn't intended that way. It just works out that way.
All of the expensive, fancy things that large corporations do lock out search engines.
If you don't have the money or the time or the knowledge to do the fancy expensive things, you are in a much better position. You are going to get much more traffic at less cost.
Large corporations doing fancy things and inadvertently locking out the search engines, then have to spend a lot of money on promotion to drive in the traffic that they threw away by doing the expensive things. Consider that logic.
First, sites that require any kind of registration or password lock out search engines.
Password or not, if there's a box you have to fill in -- this is a dumb, blind robot. It can't fill out any forms. As soon as it comes to a form, it stops. That's all the farther it goes.
Databases -- as I mentioned before, a crawler cannot get content from a database, because it cannot fill out a form.
Now if your are talking about the intranet version of AltaVista, there is for that a tool kit, which allows you to take some information from databases and put that in a form that AltaVista can index. But that works because you are in charge -- that's your database. The crawlers from the public AltaVista Search site is not going to dig into other people's databases and pull out information.
Dynamic pages -- it seems like a great idea to make it so everybody who comes to your site gets a unique experience. And there's some excellent software that let's you pull pieces out of databases to create unique user experiences based on cookies or on profile information.
But there is one problem with that. When a search engine crawler arrives, that's like dividing by zero. The crawler halts immediately, because it sees ahead of it an infinite number of pages.
This is one reason why nobody can say how many pages there are on the Web, total. Every dynamic site has an infinite number of pages. How many millions of dynamic sites are there out there?
Information inside frames -- I love this. This wasn't intentional. This is true of all search engines today. AltaVista will index what's in the outside of the frame, but not what appears in the window of the frame. Unless you have a no-frames version of your page, the information in the window of the frame does not get indexed.
I love that because I hate frames -- in most instances. There are some very good uses for frames. Actually, I was thinking of using frames for the on-line version of this tutorial, with the script on the outside, and the examples -- with live searches at AltaVista -- could appear inside the frame.
But in most instance, the frame is just a nuisance, making it easy to put a flashing ad in front of my face, and eating up my screen space. When I'm using a laptop with a small screen, it is a serious problem to have a third of my screen space eaten up with a useless frame.
One way or another, they are firmly punished by search engines for using frames.
AltaVista also won't index graphics. Have you ever been to a site that has a huge picture that takes two or three minutes to paint across your screen at modem speeds? And all the words are embedded in that .gif. A search engine can't do a thing with that. Unless the Webmaster put ALT text behind the picture, describing it and listing those important words, just like a bllind person would stop right there, the crawler stops.
Multi-media files (audio and video) and information that's in Java applets can't be indexed.
These are limitations today. I don't say that they will always be limitations. At some point in the future we will have very good voice recognition and it will become possible to index voice files. At some time, it will be possible to quickly and accurately match patterns and you may be able to search for images. But if you are designing your Web site for today, you need to be aware of the consequences of you page design for search engines.
Acrobat and PostScript files -- the intranet version of AltaVista can index those today, I believe. But the public AltaVista Search site, which drives traffic to your public Web site. the public site does not.
Another strange limitation is a pragmatic compromise, intended to help optimize the performance of AltaVista. They will only index the first 100 Kbytes of text. So if you have an entire book, it's best to break it up into chapters, and then all the text can be indexed.
They will pull out the hyperlinks from the whole document, but they will only index the first 100 Kbytes.
Comments aren't indexed at all.
When AltaVista first came out, there were a lot of wise-guy Webmasters who thought they were going to fool everybody. "I figure everybody searches for the word 'sex.' I don't have any sex at my site, but I want people to stumble across my site. So I'm going to put the word 'sex' three thousand times as comments. And any time that anybody searches for 'sex' that will come up first."
People actually tried that. They tried doing the same kind of thing in wallpaper. They tried everything imaginable to fool search engines.
But two factors come into play here. The first is that AltaVista doesn't index comments at all. The developers presumed that comments are private communications by Webmasters to themselves and to others who work technically on the site. It's information that is not meant to be public and hence should not be indexed.
The other thing is that AltaVista only counts to two. They designed that in at the the beginning. They didn't want the Web to get totally corrupted with people repeating words uselessly.
You can imagine how bad it could get. Just look at a phone book. How much of the phone book is AAAA... from people trying to be near the front of the phone book. And you can imagine how cluttered the Web would be with useless text if people thought that repetition actually got them to the head of the list. "I have 5 million copies of that word here. Why aren't I higher on the list?" It would have been an endless disaster. Fortunately, they avoided that.
My rule of thumbs is -- design for the blind. Label everything. Of course, that is easier to do with your personal site or an intranet site than it is with a public corporate site.
The blind are some of the best users of the Internet today. They use text-only browsers and text-to-voice converters, and they are able to navigate very well unless people put up these blockades and brick walls of pages with big images that you need to see to understand.
If you need to have a picture, be sure to label everything clearly with ALT text in the background, to explain what a sighted person would see.
Remember, a crawler really is very much like a blind user. Next slide
Return to B&R Samizdat Express