Dave N5UP

Dave N5UP
Dave monitoring progress during the server migration June 17, 2010

Dell R710 server

Dell R710 server
A technician at our data center adjusts a rack of Dell R710 servers

Sunday, July 18, 2010

System will be down 0400 to 0500 July 19 UTC

During the conversion to the new server farm, I failed to create one of the index files that is necessary for the Account Manager to work properly. Without this index, it takes about 30 minutes for the screen to load, which makes the program totally useless. With the index in place, it takes a few seconds.

Unfortunately, it will take about an hour to create this index file, during which time the entire system has to be taken out of service.

The Account Manager is the only tool that will allow you to split your account into multiple accounts, with different time periods or QTHes, and that will move the eQSLs around to the proper new account automatically.

We are going to bite the bullet and create the index tonight, July 18 at 11pm Central Time, which is 0400 UTC July 19. Hopefully it will take less than 1 hour, but we really have no idea how long the index creation will take on these new faster servers.

73,
Dave N5UP

Wednesday, July 7, 2010

New System is working great!

So far, the new system has been working fantastically well. With a maximum of 240 simultaneous browser connections, the application server has been running at a maximum of 10% CPU utilization. Meanwhile, the database servers is still serving over 99.5% of all database requests directly from memory instead of requiring a disk access.

Wednesday, June 30, 2010

Scary? I'll tell you what's scary... and sad...

Turning out the lights on 2 servers that have served you perfectly, without a hiccup, for 2 1/2 years, deleting all the files, hoping you did get everything moved over, checking your backup and your off-city backup, and then checking it all again, and then entering the service cancellation order and logging off for the last time.

Sad, and just a tad scary.

Friday, June 25, 2010

I am making this post from my iPhone to test the ability to post status information even if all the computers are down. This concludes the test.

The System is Up

The Application Server has been handling over 200 connections with less than 10% of the CPU resources.

The Database Server has been returning an average of 99.62% of the queries from memory without requiring a disk access.

Friday, June 18, 2010

A full day on the new servers

A full day of operation in which a couple of bugs were pointed out and fixed, but otherwise everything ran relatively smoothly. The only outstanding problem I'm aware of is that the Account Manager had to be disabled because it needs a new index, and I have to take down the entire system for an hour to build that index.

I'm looking at the performance monitor, and it shows 184 user connections only used a maximum of 9% of the CPU resource on the application server. Very cool.

The database server is the real hero here. It is able to keep almost a quarter of the entire database in memory, so that requests for data can be served from memory instead of requiring a disk access. My list of long-running queries shows the worst offenders to be running about 12 seconds. On the old servers, there were some queries that ran for 5 or 6 minutes!

Now that I can take a few minutes to actually breathe, there are 251 cards in the queue needing to be printed and mailed. Picked a bad day to run out of inkjet ink and card stock.

Thursday, June 17, 2010

First day on the new servers

One small issue: the eQSL applications were reporting the time as 1 hour ahead of the correct time that is set in Windows. I fixed that by overriding the Java Runtime Engine's default of British Summer Time instead of what I wanted, which was GMT!

Everything is going smoothly the first day on the new servers. I managed to get some sleep from 0700 to 1200 UTC.

I am adding a 1 Terabyte disk to my backup server, located in Houston Texas (so I have a catastrophic backup in the event of a disaster in Dallas). That disk should come online in a few minutes, and I will start making backups of the Application server and the Database server.

Your logger that does Real-Time uploads of ADIF logs might fail occasionally over the next 72 hours. This is because the eQSL.cc domain name may take up to 3 days to propagate all around the world. There is nothing that can be done, except to wait a day and try again. In many regions, the eQSL.cc domain is already working properly, pointing to the new servers without a change in the URL to www.QSLCard.com. It should not take more than about 72 hours for this to occur.

If you logged into www.eQSL.cc but now your URL says www.QSLCard.com you can try logging out, then point your browser again to www.eQSL.cc and log in again. If it stays logged into wwww.eQSL.cc then you are all set. If it redirects you to www.QSLCard.com then your area does not yet have the updated eQSL.cc routing. Just wait another few hours.

In any event, everybody is now on the new servers and can use the system without limitations with their browser.


The system is running much faster from what I have been able to see. Right now I am seeing over 90 users logged in, and the longest wait for a database response has been on the order of 20 seconds, compared with 10 minutes or longer on the old machine.

Even the Power Users screen, which used to take 2 or 3 minutes (if it didn't timeout first) now only takes a few seconds to display.

The new application server is easily able to handle those 90 users, and has never had to process more than 2 users at the same time, because it is handling their requests so fast. On the old machines, it was quite normal to have 10 to 20 users being processed simultaneously.

Most of the speed improvements are the result of the new database server, which has 24 Gigabytes of memory, along with 5 hard disk, of which 4 provide fully redundant, mirrored and striped (RAID 10) data storage spinning at 15,000 RPM for the database, and the 5th of which is a 1 Terabyte disk for storing database backups.

Please don't report any errors yet, unless they involve money ;) so we can have a chance to find errors ourselves and fix them. Otherwise you may overwhelm our email support volunteers with questions they cannot answer.

If everything continues to go this well, I will eliminate the time delays on the OutBox (I have already reduced it from 10 minutes to 1 minute) and the InBox, and other screens.

If you have questions or comments, feel free to post them here on this blog.

73!
Dave