Google, Bing & Beyond: How the Internet Finds You

The 80:20 principle

Every time you “Google” something, you’re using a search engine. The World Wide Web – the internet – is a network of interconnected pages and resources, accessible to users on a global scale. To find the web page, and you know its address, the URL (its uniform resource locator) you can just go straight there. Most of us, though, type just a word or phrase into a search engine. These keywords and search terms and the results that search engines deliver, present a priceless opportunity for business owners who want users to find their goods and services. 

Because most people, especially digital natives, take this for granted, most people don’t pay much attention to what it takes to make search engines do what they do. We’re going to unpack, hopefully without too much tech-speak, how search engines work, and why understanding how they work, is vital to your business.

Get coffee. In trying to cut through the technospeak, this is quite a long read.

Lots of search engines but just one dominates

Google wasn’t the original and isn’t the only search engine around. It’s pretty obvious why “Google” has replaced the verb “search” in everyday language.

Google has the largest market share followed by Bing. It’s easy to see why: Google had a huge head start, launched in 1998, and Bing arrived 11 years later in 2009. Even backed by the resources of Microsoft, it hasn’t been able to catch up. Third in line is the “old man” on the block, Yahoo!, which like Google and Facebook, had its genesis on a university campus back in 1994. Unlike Google and Bing, which began as dedicated search engines, Yahoo!, virtually from the get-go offered a suite of other services like e-mail and news in addition to its directory. It is this directory that evolved into one of the earliest search engines.

Going back in time to get to the future present

Some of us (like this writer) remember the days before the Internet. We also remember the days when (re)searching (for) something meant visiting your local library and/or consulting an encyclopaedia. To find your way around the encyclopaedia, you would consult its index to find a list of the volumes and/or pages of information about your selected subject. 

Outside of research for a school or academic project, everyday research, like checking the prices of goods and services, meant either a series of telephone calls or crawling your way around the shops – in person. If you were lucky, you might also be armed with information you’d collected from the local newspaper, advertising and word-of-mouth recommendations.

Back in the library, you’d start your search by going to the catalogue which consisted of a series of index cards in banks of drawers. The catalogue recorded each and every book in the library, organised into subject groups and sub-groups or themes using a system devised by Melvil Dewey in 1873. By flipping through the catalogue, you could identify which books in the library dealt with your subject and then begin to refine that search to identify which books would best suit your specific research purposes.

As an aside: libraries still use the Dewey decimal system for organising themselves and their books. 

Catalogue vs index: same but different? No

A catalogue, whether electronic or hard copy, is simply a record of things that have been classified and described. For example, if you shop online and you’re looking for a new device, you check out a vendor’s website and look for the brand, model and price that best suits your pocket. You’re browsing the vendor’s online catalogue. 

An index, however, works differently. Even though an index usually refers to a single book, once you look up your subject, under the main topic, you’ll find a series of sub-themes. For example, if you have a recipe book, you may want to know how many different ways you can use a particular ingredient, like cheese. The index will point you to different sections of the book where cheese – and different types of cheese – can be used, e.g., starters, main courses, desserts or vegetarian dishes. Then there’ll be another “sub” list of specific recipes, for example cheesecake, cheese scones or Quiche Lorraine. 

In summary: a catalogue is a list while an index is a highly organised list with classifications and links between subjects and themes.

What do old fashioned library catalogues and indices have to do with search engines?

Thank you for asking: indexes and catalogues have everything to do with search engines. For search engines to render – deliver – results on your search, the sources of that information must be found, catalogued (organised and described), indexed (broken down into sub- and linked categories) and ranked. 

Dedicated web crawlers

The World Wide Web is enormous and it’s getting bigger all the time as new content is constantly generated and added. For search engines to find the results you want, the companies behind them have bots that crawl the internet – like spiders would their own webs – identifying keywords (tokens), downloading pages, extracting information and identifying links to other pages. They use this data to work out whether sites have new pages or if content on existing pages has changed. The links also help the web crawler to discover new pages to add to its ever-growing database.

Each search engine has its own web crawler and webmasters know that their sites have been visited when the search engine leaves its “calling card” – a user agent string like the ones below:

  • Googlebot User Agent
    Mozilla/5.0 (compatible; Googlebot/2.1; +https://www.google.com/bot.html)
  • Bingbot User Agent
    Mozilla/5.0 (compatible; bingbot/2.0; +https://www.bing.com/bingbot.htm)

What about images and graphics?

So far, we have focused on text-based content. Because this image and video content isn’t readable to web crawlers, they have to “look” elsewhere, like in the file name and metadata (also known as the alt text) for the information they need. This explains why web designers and webmasters will get tetchy with clients about image file names and descriptions.

From catalogue to index

Clearly, it’s not enough to identify and create a mammoth list of the world’s websites and their pages. This list needs to be organised in a way that when we do an internet search, we get the instant results we expect. Remember we mentioned keywords or tokens? It’s these keywords, among other things, that the search engine uses to organise the websites in its database so that our searches yield optimal results.

Ranking: the magic and mystery of the algorithm

Once the data – web pages – have been collected and organised, they need to be ordered so users’ search results make sense. In other words, the results the search engine renders are, in fact relevant and useful. This is the mysterious magic of search engines that is so important for business owners to understand. 

To do their jobs, i.e., rank websites, search engines need to answer a series of questions:

  1. What does the user – searcher – mean? Why have they chosen those words? What lies behind their search? 
  2. What pages are relevant to that question?
  3. Is there quality content?
    Criteria for identifying sites with superior content include, among other things, the number and quality of any (back) links within the site and to other credible sites.
  4. Is the page technically useable?
    In other words, does the site load quickly, is it secure, etc.

To solve this conundrum, search engines use algorithms. We all often refer to “the algorithm” without really understanding what it is. In its purest definition, an algorithm is the procedure for solving a specific mathematical problem. More broadly, an algorithm is a step-by-step process for getting to a solution which is what search engines do for their users. Finding out how a search engine’s algorithm works is virtually impossible. 

A search engine algorithm will often be a trade secret, so search engine companies keep that information confidential. Also, algorithms change based on both their owners’ objectives and the nuances of the rapidly transforming World Wide Web.

Current, quality content is king

Like with so many things in life, the 80:20 principle applies to aspects of the search engine world. The bots and crawlers do 80% of the behind-the-scenes work with the algorithm doing the all-important and detailed connections necessary to render the results users need. 

But that’s not enough. If business owners want users to find their websites and pages – in other words, for their websites to rank – they must pay attention to the content and structure of their sites. 

Search engine optimisation – more than words

Carefully crafted content – and carefully selected keywords – are only part of what SEO is all about. As we noted earlier, content is only one of the four criteria that search engines use to rank websites. The content needs to be current (regularly updated), interesting and well-written and include links to, and from, other sites. Then there are the typical technical considerations of the website, itself, which among other things, should be easy to navigate.

We get how search engines work

At Polar Bear, we get how search engines work. For us, SEO is not just about the words, it’s about optimising every part of your company website so that it becomes an integral digital asset in your broader marketing arsenal.