Sun has identified some issues with our disk array configuration. They have provided some new settings for us to apply. One of the changes has been made, with little effect on disk performance. The other two changes will require that the email server be shutdown. We are scheduling this right now. Please refer to the “Scheduled Outages” (http://www2.ups.edu/ois/nssg/network/alerts.shtml)page for the latest information.
IMAP and POP3 and Webmail Slowness I
We are currently experiencing disk performance problems on the mail server. This is causing slowness and problems connecting with Webmail and POP3 and IMAP clients. We are working with the hardware vendor to correct the problem.
WeMail Problems I
Today at about 10:00 AM, the HelpDesk reported major failure in WebMail.
We noticed the presence of a large number of processes (about 1000 and growing) on the mail server, and a larger than normal of mail in-queue. WebMail was stopped, server processes on the mail server were stopped, and the mail queues were processed by hand.
Continue reading
SAN Maintenance and New Installations
On Saturday we will restructure the fibre channel fabric by replacing the existing 8 port switch with two 24 port switches—32 total ports enabled. As part of the restructuring, the database systems (Lenel, Rainier, Crystal, and Grace) will be down. The SAN and tape library will also be down during the restructuring.
We will be attempting a variety of hardware related tasks during the down time:
1. Turning the two Dell cabinets 180 degrees to provide better air flow.
2. Reworking the network, fibre, and power cables in the Dell cabinets for better access.
3. Migrating and rezoning all fibre channel connections to the new switches.
4. Installing a new HBA in Rainier and upgrading the PowerPath license.
5. Adding additional systems to the fibre channel fabric (Exchange Servers, AX100 appliance, Veronica, Merlin2 and Alexandria).
We estimate that this will take approximately eight hours. We will be starting at 8:00 am on Saturday, June 11
Change to PureMessage
Over the last few weeks a pattern has appeared in the cpu load on the sophos server: over the course of time the cpu load increases and stays high (
PureMessage not accepting messages
PureMessage stopped accepting messages this morning when the disk volume was filled by log files. Corrections have been made to the logrotate.conf file in an attempt to prevent this from occurring in the future.
modification to mx records
MX record added for second interface on gehenna to reduce mail load to sophos. Entry created in firewall and external dns for anneheg.
Stray Entry Hit with Comment SPAM
One of our Maintenance Blog entries was accidentally left open to comments, and was filled with “Comment SPAM”. The offending entry was reconfigured, and the SPAM was deleted.
E-mail problem: 5550 5.3.0 Can’t create output
Some user reported this morning the inability to send messages to users. They received a common error,
Final-Recipient: RFC822; username@ups.edu
X-Actual-Recipient: RFC822; username@ups.edu
Action: failed
Status: 5.3.0
(reason: Can’t create output)
This error was the result of poor deactivation of the quotaing system. The quota system had been turned off, but not removed from the fstab file. When the system was rebooted yesterday, the quota system was re-enabled and locked account in excess of their time limit. This issue has been corrected.
Network upgrade
The planned OS upgrade of core network equipment on Sunday was not as smooth as planned. Two systems had difficulty with the upgrade and required a reboot: the mail server, and the Oracle development server. Otherwise, the upgrade was a success.
As a result of the problem, we have a better understanding of the upgrade process for future upgrades.
Added database to mySQL
Added new database and user to MySQL for testing purposes.
Added a redirect in /.htaccess
Added to /.htaccess in www.ups.edu
Redirect /sciencecenter http://sciencecenter.ups.edu/
This is in support of an awareness campaign in which the address www.ups.edu/sciencecenter is in use. This will allow the includes to work at all levels without a major reworking of the site.
Packeteer Fails
The Packeteer failed on 12/15/04. In fact, it’s been failing for the past few months with a hard drive problem, but today, it went down hard. Packeteer is shipping a new unit to arrive 12/17/04. Until then, the traffic on our internet is uncontrolled.
Dial-in Failures
The radius daemon and portmaster were reset to address several reported issues with dial-in access. The core reason for users not being authenticated is not completely clear, but there are indications that the portmaster or the radius daemon became confused about the appropriate share secret. Once all entryies were reset authentication started to be validated correctly.
During the process, four modems were identified as failing to respond properly and were removed from service.
Web Server is Back!
The University web server has been repaired. All functions are operational, including ftp access have been resored. Full service was restored at 7:30 PM.
Continue reading