Saturday, August 19, 2006

Richard Skinner's groundbreaking book "IT is about the Strategy" is a comprehensive guide to developing winning IT strategies for the SMB.

Most young organizations suffer from the same breakage points as they grow. Skinner shows how these milestones and there accompanying issues are a normal part of business growth, and how you can overcome each of these challenges.

I have had the pleasure of working with Richard a number of times over the years, and I know these techniques work. His book outlines common sense techniques that will take you out of firefighting and chaos. If you are looking for practical help with your IT issues, please check out this book at www.itisaboutstrategy.comLink

Monday, August 14, 2006

The Daily Status Meeting

In most businesses we face IT issues on a daily basis, dealing with them in a timely, and effective manner is critical. I have found it extremely useful to meet with the IT staff each morning and review any of the issues currently open or, just recently closed in our tracking system.

We limit this meeting to just the critical service impacting events. Having this meeting early in the morning gives us a chance to get the necessary resources deployed to deal with the problem effectively. Including all of the IT disciplines is essential, make sure that there is a good discussion of the issue and review how process is being followed in updating tickets, and communicating with the users.This is perfect opportunity to drive the internal IT processes and get everyone on the same page.

Many of the things I do today have been gleaned from the morning meeting. Remember, we are always on the journey to better IT management, we however never arrive. Maintaining effective processes including incident and problem management is a continuous effort. When IT staff are not consistently directed they will fall back into old habits of poor documentation, and poor communication.

I also use this meeting to review any changes that went into production the previous day. This is the perfect time to check status on these items and it only takes a few minutes. A well run morning status meeting will typically take 15 to 30 minutes, certainly a small price to pay to know what is going in your IT environment.

This may seem overly simplistic, but you would be amazed at how many organizations do not do this simple task.

Sunday, August 06, 2006

Are there babies dying?

With the recent turmoil in the middle east, there is indeed a situation where babies are dying. It is tragic to watch the news and see so many innocent children caught up in this tragic violence. Despite varying political views, no one wants to see young children suffer. We see rescue workers on both sides struggling to give aide to these victims as they place there own lives at risk. This is no doubt an emergency situation. The point of this post is to contrast these "real life" emergencies with the IT emergencies that we deal with.

On countless occasions I have been drawn into panic situations on the IT front, where it appeared as if certain doom was imminent, if the server wasn't back up right now. Managers are screaming, the phone rings off the hook, vacations are called off, and the world is on the edge of destruction, right? I worked with gentleman a few years ago, and while in the middle of one of these IT catastrophes asked "Are there babies dying?". This really hit home with me and helped me put things in the proper perspective. Large revenue impacting events are definitely important and should be responded to appropriately, but they are not on the same scale as the life and death events we see around us. Organizations find themselves in a panic first mentality which is counter productive, stressfull, and not cost effective.


In dealing with the IT emergency, some well thought out procedures can deal with the most critical events imaginable. The organization I work for had several incidents with hurricanes last year which required the shutdown of offices, rerouting of business, and the human task of getting money to people while the office was out of commission. We were able to do this, without undue stress, efficiently, and with some compassion. All of this was possible because we took time to plan, when the emergency arose, we new what to do.

Here are some really basic tips on disaster planning that we often take for granted.

1. Maintain an up to date contact list, with roles and responsibilities. It is amzing durng a crisis how hard phone numbers are to find.

2. Have a communication plan. Know who communicates with clients and who will manage internal communication. Let your staff do there jobs and have someone else manage communication. Know what you will do if all Internet, and telephone communication is shut off.

3. Prepare an escalation process, and make sure everyone knows how to use it.

4. Keep up to date inventories of equipment and software.

5. Backups, backups, backups. Enough said.

6. Document the chain of events. Use your incident management system if available to keep track of all activities surrounding the emergency.

7. At the end of the event, do a postmortem. Review the time line of events and look for things that could have been done better. Be critical of yourself, this will pay huge dividends the next time an event occurs. I found that sharing the weakness in our responses and a solid plan to correct them next time, adds credibility with clients and auditors. Don't be afraid of the truth!

In conclusion I would like to say that some cool heads, and basic management techniques solve problems faster, and more effectively than panic, and finger pointing. Remember the next time a server crashes ask "Are there babies dying?".

Wednesday, July 26, 2006

My server just crashed, now what???

The phone rings and on the other end is an angry user complaining that the system is down again. This is every IT persons nightmare, not to mention the user.

The IT person then scurrys around making phone calls, sending out frantic pings, hoping that the problem isn't in there area. System admins blame the network, network says there isn't a problem on there end, and send the problem back to the system admin, who in turn blames the programmers. What do you do? You reboot the box, when the box comes up the sys admins say they can't find anything in the logs, and it must just be one of those things. This happens about once or twice every 60 to 90 days. Does this sound familiar?

A system reported as down or crashed can mean many different things to IT, but to the user the system they depend on is down, end of story. Sometime systems crash or lock up and there is no apparent reason readily available, most of the time however there are warning signs and the mature IT staff will be looking for them.

This isn't an article on system administration, my point here is about monitoring what goes on in your environment. I shudder every time I here of a system locking up becuase of lack of disk space, 90% of the time this problem can be avoided by some simple monitoring and alerting. Monitoring will assist with other issues such as lack of memory, cpu, network traffic, or other errors. There is no excuse for organizations not doing network monitoring, There are free tools available. Nagios for example can be configured on a fairly low end machine and can track virtually anything that needs to be monitored. Nagios alows for alerts to be generated, escalations done, and has some rudimentary reporting which can show trending of uptime.


Having data available will help pin point problems much quicker. In the event of a server down scenario, a quick scan tell you what the network is doing, how the server was performing, and what if any capacity issues exist. With the setting of proper thresholds, most problems can be avoided before the user experiences a problem.

Most small businesses don't have good monitoring in place. If there is monitoring, it is done by different groups for there own particular needs, and a comprehensive approach is not in place to give everyone access to this vital information. Why is this? Poor managment. It is a management responsibility to make sure this infomation is available and regularly reported on.

Don't let your systems go unmonitored. Check out www.nagios.org this tool could save you thousands of dollars in lost production time. There are many other useful tools available, Nagios is the one I am most familar with, and in my opinion offers the most versatility. If you need help configuring Nagios, shoot me an email and I will help get you to a resource who can solve your problem.

Wednesday, July 19, 2006

ProjectSteps: A Collection of Project Management Sayings

Blogger Stephen Seay has assembled a collection of project management sayings that capture the thrill of victory and the agony of defeat experienced by the IT project manager.

Click on this link
ProjectSteps: A Collection of Project Management Sayings

What's in a project list?

I am always amazed when I go into a new company and ask to see the list of IT projects. Typically a spreadsheet gets produced with a wide variety of items that are typically 1-2 years old. When you talk to the admins, and programmers, they more often than not, have never seen "the project list".

Most small IT organizations have a huge back log of work that never gets done, the project list is more of a wish list! If you are going to move the organization along and get some of this strategic stuff done, you have to stay focused. Break these big projects down into manageable chunks and hold people accountable. Decide what is important and continue to drive forward, remembering that you can always find a small enough piece that can actually get done.

The other thing that amazes me is that many organizations have projects that make up a significant portion of there budget and there has never been a project definition, or scope defined. The project is underway, resources are being utilized, and a firm deliverable is nowhere in sight! If you have any of these "projects in the wild" stop right now.

While most small IT shops don't have budget money to spare, they continue to put resources into projects that will never come to fruition. I have seen several companies spend literally millions of dollars on the "system consolidation" project which will get rid of all of there core legacy systems and replace them with a new all in one system developed from scratch. I have yet to see one of these come into production.

Remember ...

Keep it simple.

Hold people accountable.

Identify the requirements, don't get caught in the "Ready, Shoot, Aim" game.

If you have a horror story of a "project in the wild", let me know.