PhpDig excels at small Web site indexing

Submitted by srlinuxx on Thursday 8th of December 2005 11:25:56 AM Filed under

HowTos

Webmasters looking to provide search capabilities for their site would do well to try out PhpDig, a Web spider and search engine written in PHP with a MySQL backend. There are other open source search engines, all of which have their own advantages. PhpDig just happens to suit the needs of my Information Technology for Greenhouses and Horticulture site. Here's how I got it working.

Webmasters with small sites know the problem of providing useful site search capabilities. Typically, visitors enter keywords in a search box and the search engine returns a ranked list of pages related to the query. This is a useful service -- provided the visitor can tune the search and the results returned are reliable and relevant.

Some Webmasters rely on Google for this service. A listing in Google or another mainstream engine is a must-have in practical terms, so it is easy enough to piggyback on the main engine with a site-specific search, provided Google understands your site and keeps coming back for updates -- but this isn't always the case.

Large search engines boast of indexing of billions of pages, but we are only interested in digesting a hundred pages or so. We need them indexed on a regular basis, daily or at least more often than Google might do it.

It is also important to know if our site is responding correctly by providing public pages, hiding private pages, and following links correctly. Since Google uses algorithms that it doesn't share, we have no way of predicting the indexing results or doing any testing in advance. Advance testing is useful if, for example, you have private files that you want to be sure will not be indexed, but you are relying on your robots.txt file to deny access to bots. If we make a spelling mistake in robots.txt, our private pages could go in Google's cache for the world to read. We also need to control what words are indexed and customize our own search and result pages.

Enter PhpDig.

Login or register to post comments
Printer-friendly version
2005 reads
PDF version

More in Tux Machines

digiKam 7.7.0 is released

After three months of active maintenance and another bug triage, the digiKam team is proud to present version 7.7.0 of its open source digital photo manager. See below the list of most important features coming with this release.

Dilution and Misuse of the "Linux" Brand

Linux Foundation Rewards StepSecurity’s Impact on CI/CD Pipeline Security Fixes for Critical Open Source Projects [Ed: Having just participated in a FUD attack together with a Microsoft proxy, not to mention issued a report with it]
Cardano Roundup: Lace Wallet Announcement, Hoskinson Proposes Self-Regulation, and Linux Foundation Membership [Ed: The "Linux" Foundation misuses or sells the Linux brand, diluting the name and the project's identity]
Can SONiC be the Linux of Networking? [Ed: The Register now abuses the Linux brand to describe something of Microsoft, which is attacking Linux]
Kuro: An Unofficial Microsoft To-Do Desktop Client

Microsoft says that they love Linux and open-source, but we still do not have native support for a lot of its products on Linux.

Samsung, Red Hat to Work on Linux Drivers for Future Tech

The metaverse is expected to uproot system design as we know it, and Samsung is one of many hardware vendors re-imagining data center infrastructure in preparation for a parallel 3D world. Samsung is working on new memory technologies that provide faster bandwidth inside hardware for data to travel between CPUs, storage and other computing resources. The company also announced it was partnering with Red Hat to ensure these technologies have Linux compatibility.

today's howtos

How to install go1.19beta on Ubuntu 22.04 – NextGenTips

In this tutorial, we are going to explore how to install go on Ubuntu 22.04 Golang is an open-source programming language that is easy to learn and use. It is built-in concurrency and has a robust standard library. It is reliable, builds fast, and efficient software that scales fast. Its concurrency mechanisms make it easy to write programs that get the most out of multicore and networked machines, while its novel-type systems enable flexible and modular program constructions. Go compiles quickly to machine code and has the convenience of garbage collection and the power of run-time reflection. In this guide, we are going to learn how to install golang 1.19beta on Ubuntu 22.04. Go 1.19beta1 is not yet released. There is so much work in progress with all the documentation.
molecule test: failed to connect to bus in systemd container - openQA bites

Ansible Molecule is a project to help you test your ansible roles. I’m using molecule for automatically testing the ansible roles of geekoops.
How To Install MongoDB on AlmaLinux 9 - idroot

In this tutorial, we will show you how to install MongoDB on AlmaLinux 9. For those of you who didn’t know, MongoDB is a high-performance, highly scalable document-oriented NoSQL database. Unlike in SQL databases where data is stored in rows and columns inside tables, in MongoDB, data is structured in JSON-like format inside records which are referred to as documents. The open-source attribute of MongoDB as a database software makes it an ideal candidate for almost any database-related project. This article assumes you have at least basic knowledge of Linux, know how to use the shell, and most importantly, you host your site on your own VPS. The installation is quite simple and assumes you are running in the root account, if not you may need to add ‘sudo‘ to the commands to get root privileges. I will show you the step-by-step installation of the MongoDB NoSQL database on AlmaLinux 9. You can follow the same instructions for CentOS and Rocky Linux.
An introduction (and how-to) to Plugin Loader for the Steam Deck. - Invidious
Self-host a Ghost Blog With Traefik

Ghost is a very popular open-source content management system. Started as an alternative to WordPress and it went on to become an alternative to Substack by focusing on membership and newsletter. The creators of Ghost offer managed Pro hosting but it may not fit everyone's budget. Alternatively, you can self-host it on your own cloud servers. On Linux handbook, we already have a guide on deploying Ghost with Docker in a reverse proxy setup. Instead of Ngnix reverse proxy, you can also use another software called Traefik with Docker. It is a popular open-source cloud-native application proxy, API Gateway, Edge-router, and more. I use Traefik to secure my websites using an SSL certificate obtained from Let's Encrypt. Once deployed, Traefik can automatically manage your certificates and their renewals. In this tutorial, I'll share the necessary steps for deploying a Ghost blog with Docker and Traefik.

Best Features in Ubuntu 23.04 “Lunar Lobster”

Ubuntu 23.04, codenamed Lunar Lobster, is the latest version of the popular open-source operating system to be released on 20th April. It comes with a host of new features and upgrades designed to enhance the user experience and improve the system’s overall performance.

Top 7 Linux Window Managers

Linux is known for its open-source nature and flexibility. One of the essential components of a Linux system is the window manager. A window manager is responsible for the appearance and management of windows on the screen.

Navigation

Language Selection

Search

Tux Machines Section/s

Authors Information

LWN

Active forum topics

Collabora

9to5Linux

Gnu Planet

FOSS Force

FOSSLinux

Phoronix

OpenSource.com

Kde Planet

Fedora Magazine

Mozilla

PhpDig excels at small Web site indexing

More in Tux Machines

digiKam 7.7.0 is released

Dilution and Misuse of the "Linux" Brand

Linux Foundation Rewards StepSecurity’s Impact on CI/CD Pipeline Security Fixes for Critical Open Source Projects [Ed: Having just participated in a FUD attack together with a Microsoft proxy, not to mention issued a report with it]

Cardano Roundup: Lace Wallet Announcement, Hoskinson Proposes Self-Regulation, and Linux Foundation Membership [Ed: The "Linux" Foundation misuses or sells the Linux brand, diluting the name and the project's identity]

Can SONiC be the Linux of Networking? [Ed: The Register now abuses the Linux brand to describe something of Microsoft, which is attacking Linux]

Kuro: An Unofficial Microsoft To-Do Desktop Client

Samsung, Red Hat to Work on Linux Drivers for Future Tech

today's howtos

How to install go1.19beta on Ubuntu 22.04 – NextGenTips

molecule test: failed to connect to bus in systemd container - openQA bites

How To Install MongoDB on AlmaLinux 9 - idroot

An introduction (and how-to) to Plugin Loader for the Steam Deck. - Invidious

Self-host a Ghost Blog With Traefik

Latest News

Best Features in Ubuntu 23.04 “Lunar Lobster”

Top 7 Linux Window Managers

Tux Machines IRC Logs 2021 Archive

2021

Syndication & Social Media

Support Tux Machines

Recent comments

Who's online

Gizmoz

FOSSMint

Debug Point

LXer

UbuntuBuzz

SparkyLinux

Linux Journal

Linux Today

System 76

Kernel Planet

Purism

LinuxSecurity.com Advisories

Debian

It's FOSS

OMG! Ubuntu!

GamingOnLinux

Slashdot

DW Latest Releases

DistroWatch

FSF

Ubuntu Buzz

LinuxLinks