Dave N5UP

Dave N5UP
Dave monitoring progress during the server migration June 17, 2010

Dell R710 server

Dell R710 server
A technician at our data center adjusts a rack of Dell R710 servers

Wednesday, October 15, 2014

Server Outage October 15, 2014 The database server was migrated to another rack. Everything is OK again.

Tuesday, April 10, 2012

Some server maintenance the week of April 10-13

Well, I had planned to migrate the system over to a new server over the course of this week, methodically and carefully. But all that changed this morning in a matter of minutes when I was rebooting the system after a Windows Update, and the server failed to boot up.

I have thus moved the domains eQSL.cc, eQSL.org, and eQSL.net over to the new server, and it may take a little while for them to all point to the new server. If your favorite domain doesn't go there yet, you can try one of the others. A fourth domain, eQSL.CO has been pointed there for some time already, and is guaranteed to work.

The problem with not doing an orderly and methodical migration is that any user graphics uploaded after February 16 are not on the new system. If you uploaded your license for Authenticity Guaranteed, and you still have not received it, your license image is also lost, so you should upload it again.

Over the next few days, I should be able to synchronize all the graphics files, so that your eQSL graphics are all up-to-date. If you don't want to wait, you can always upload them again right now. I will not overwrite newer files.

Monday, March 5, 2012

Server Outage March 5 2012

The system is back up and running normally again.

We experienced a server outage at around 0930 UTC. You may have seen error messages to the effect of a transaction log being full.

The system was back up at 1100 UTC.

Thursday, January 19, 2012

We are rebooting a server

The system is down for a few minutes while we reboot a server that is misbehaving.

Monday, December 12, 2011

System was down temporarily on Monday

The database server failed to make a transaction log backup. And we had a glitch in the notification system, so it took us longer than usual to notice. But we got everything working in about hour after we discovered what had happened, and the system is back to normal now. Sorry for the outage!

Sunday, December 4, 2011

Very short system outage

We are rebooting a server, so the system will be down for a few minutes

Monday, October 24, 2011

Database outage - Fixed

The database outage that began around 5:30am Central Time is now fixed. One of the log files in our database had overflowed and was creating a system outage. We have freed up the necessary disk space, and rebuilt the log.

Sunday, April 10, 2011

We are rebooting the server

The application server hung up, so we are taking this opportunity to install a bunch of Windows updates and reboot it. It should be back up by 1515 UTC at the latest.

Wednesday, January 19, 2011

Possible system delays or short outages Thursday

The network operations center has informed us that some network routers need to be replaced sometime Thursday. You may experience a short outage or system delays while the work is being done.

There is no need to contact us, as we will be very aware of any outage, and the system should return to service quickly if it does go down for a few minutes.

Monday, January 3, 2011

Server outage 0630 UTC today - Fixed

We had a problem with one of our backups running out of disk space, which temporarily halted database activity. It has been fixed.

The backup file was over 300 gigabytes in size. We received the alert at 0622 UTC and immediately found the problem. We freed up enough disk space to restart the backup, and it finished without any problems.

The system was running again at 0745 UTC.

Tuesday, September 21, 2010

Server outage 0823 UTC Tuesday - Fixed

As of 13:40 we are back up and running again.

We had been experiencing some delays and lockups on a server, so I applied some new Windows and ColdFusion updates and restarted the server. For some reason it did not come back online. There were 5 different data center personnel involved in troubleshooting and fixing the problem. The primary disk passed all the tests. The RAM was replaced, and the BIOS was updated. Finally we ran CHKDSK on the drive repairing any errors. At 1300 UTC we reinstalled the operating system, and that finally fixed the problem.

Thursday, September 16, 2010

System is up again

The problem with the network router at our hosting facility has been fixed, and the system should be accessible to everyone again.

Network router has made the system unavailable

Our hosting facility has had a network router failure. Our servers are behind this router, and so you cannot access our servers until they fix the router.

Sunday, August 1, 2010

Up-time Stats for July 2010

From our Fort Lauderdale, Florida testing center:
July 1 to August 1, 2010

2,973 attempts,
2,973 successful hits,
100.00% up-time
Average response time: 3.33 seconds for page load

Sunday, July 18, 2010

System will be down 0400 to 0500 July 19 UTC

During the conversion to the new server farm, I failed to create one of the index files that is necessary for the Account Manager to work properly. Without this index, it takes about 30 minutes for the screen to load, which makes the program totally useless. With the index in place, it takes a few seconds.

Unfortunately, it will take about an hour to create this index file, during which time the entire system has to be taken out of service.

The Account Manager is the only tool that will allow you to split your account into multiple accounts, with different time periods or QTHes, and that will move the eQSLs around to the proper new account automatically.

We are going to bite the bullet and create the index tonight, July 18 at 11pm Central Time, which is 0400 UTC July 19. Hopefully it will take less than 1 hour, but we really have no idea how long the index creation will take on these new faster servers.

73,
Dave N5UP

Wednesday, July 7, 2010

New System is working great!

So far, the new system has been working fantastically well. With a maximum of 240 simultaneous browser connections, the application server has been running at a maximum of 10% CPU utilization. Meanwhile, the database servers is still serving over 99.5% of all database requests directly from memory instead of requiring a disk access.

Wednesday, June 30, 2010

Scary? I'll tell you what's scary... and sad...

Turning out the lights on 2 servers that have served you perfectly, without a hiccup, for 2 1/2 years, deleting all the files, hoping you did get everything moved over, checking your backup and your off-city backup, and then checking it all again, and then entering the service cancellation order and logging off for the last time.

Sad, and just a tad scary.

Friday, June 25, 2010

I am making this post from my iPhone to test the ability to post status information even if all the computers are down. This concludes the test.

The System is Up

The Application Server has been handling over 200 connections with less than 10% of the CPU resources.

The Database Server has been returning an average of 99.62% of the queries from memory without requiring a disk access.

Friday, June 18, 2010

A full day on the new servers

A full day of operation in which a couple of bugs were pointed out and fixed, but otherwise everything ran relatively smoothly. The only outstanding problem I'm aware of is that the Account Manager had to be disabled because it needs a new index, and I have to take down the entire system for an hour to build that index.

I'm looking at the performance monitor, and it shows 184 user connections only used a maximum of 9% of the CPU resource on the application server. Very cool.

The database server is the real hero here. It is able to keep almost a quarter of the entire database in memory, so that requests for data can be served from memory instead of requiring a disk access. My list of long-running queries shows the worst offenders to be running about 12 seconds. On the old servers, there were some queries that ran for 5 or 6 minutes!

Now that I can take a few minutes to actually breathe, there are 251 cards in the queue needing to be printed and mailed. Picked a bad day to run out of inkjet ink and card stock.