Posts

Part 2 - Disaster Recovery with SRM and vSphere Replication

Image
In the previous article we went through the installation and configuration of the SRM and vSphere infrastructure. The time has now come to actually doing some tests and failover some VM's.  In the simple scenario I expect everything to go smoothly but there are a few things that I'm concerned about at this point since the protected environment I'm ultimately is failing over isn't so simple. It has multiple dvSwitches, vLans, load balancer, firewall, ldap, a set of test drivers to verify the integrity of the system and access to the system in a sensible way for administrators so there are a few things that needs to be resolved. Protected Setup Test Failover The small test In this scenario I do a test failover of a single machine and verifies that it starts and that I can log into it. The steps are: Setup the machine running RHEL 5.5.  Install VMware tools  Kick of the vSphere Replication Create a Protection Group and ... no that didn...

Part 3 - Disaster Recovery with SRM and vSphere Replication

Image
In the previous articles I have described the setup of SRM and vSphere Replication as well as configuring a small and simple recovery plan (one single machine). Bursting the test bubble when doing a test failover though an additional machine.   In this part the focus will be on the more complex situation when you have a system over multiple VLANs that you want to failover. In the protected site load balancing and routing is done by the load balancer and for multicast by the firewall. Unfortunately these are still HW appliances in my case and I don't feel like running to the datacenter and move them each time someone does a test failover :-) So we need to figure out a way to provide that functionality in the test bubble. Who knows in the end the network guys might actually be willing to discuss virtual load balanced and firewalls. Protected Setup To keep it simple let us only consider two Vlan's in the Protected Setup. Vlan01 and Vlan50. VM01 communic...

Part 1 - Disaster Recovery with SRM and vSphere Replication

Image
Abstract  The company I work for has needs for disaster recovery but there is also external demands for it. The company not being very big has so far not been able to implement any automated disaster recovery solution although we do have a documented disaster recovery plan. The plan regardless of how good or bad it is isn't tested in a long time and the RTO (Recovery Time Objective) is wished by management to be days but I can't see that it is anything else than weeks. Since we virtualized our production system I've been looking at VMware vCenter Site Recovery Manager as a driver but the cost array based replication has stopped any attempts dead in their tracks. VMware vSphere® Replication has been around for a while now and we hope that it's hardened enough for a slightly bigger production critical implementation. In this article I will try to document and explain how we experiment, design and implement disaster recovery and even more important, in my opinion...

Intermittent and excessive loss of UDP traffic

This is a call for help to get an understanding of what happened. We are still flabbergasted and confused. Lets start from the beginning. We have s system consisting of a number of servers that primarily is communicating with each other through TCP. In very few occasions we use UDP. The system is running in a virtual environment (VMware). From one of the services, lets call it US (Udp Sender) we send UDP to another server UR (Udp Receiver).  The US is one or more and the UR is behind a load balancer even though it currently is a single server. After upgrading VMware from 5.0 to 5.1 the UDP packages on the UR end are lost excessively and intermittent. We are using RHEL 5.5 in a quite old version as the guest OS. Checking with tcpdump on the US side we see that all servers send packages as expected. Through divide and concur we rule out switches, load balancers, firewall, physical NICs until we on have the virtual NIC left. So far so good - now for the stra...

I absolutely HATE *nix memory metrics

Today was yet another day at the office and large scale PANIC erupted. Managers running around as if the end of the world was just around the corner. All the fuzz came from the interpretation of free in our production system, the server had only 85 MB free memory left and the end was near!  When taking a closer look it was obvious that the PANIC was unmotivated and that the server had 6.3 GB "free".  Still every time it has been a few weeks since I looked at the metrics I have to go through the following mental process to sort it out: "CRAP!!!  The end is near ... wait! ... calm down .. its linux ... its the  second line that counts.   ... so ... this .. is .. good?! .... sigh of  relief "  Thus I think that its in its place to repeat the memory metrics, what they mean and how to interpret them in yet another blog entry. free -m total used free shared buffers cached Mem: 5963 5581 3...

Possible SYN flooding on port 3306 (MySQL)

The system setup is such that the MySQL servers that put the "Possible SYN flooding on port 3306" in the log files only are exposed to system internal backend services. These in turn aren't exposed to the wild wild web. Fronting the system we have the servers publishing services to internet. Thus I was kind of stunned when the log messages started to appear and even though we had done a release of the system I found it far fetched that we should start to DOS our self. So why did the messages appear? Two different error messages could be identified in the log files and they seem to be related, especially since the Java servers with link failure do communicate with the MySQL servers. Possible SYN flooding on port 3306 @ MySQL server Communication link failure @ Java backend servers After some tcpdumping, head scratching and googling I think I have it down to the root cause, hidden in how TCP works in general, and the OS config in combi...

Organizational Patterns of Agile Software Development

James o. Coplien and Niel B. Harrison Community of Trust It is essential that people in a ateam trust each other; otherwise, it will be difficult yo get anything done. Do things that explecity demonstrate trust. Managers, for example, should make it overtly obvious that they facilitate the achievement of organizational goals, rather than playing a central role to assert control over people. Take visible actions to give control over process. Both overtly ambitious schedules and overly generous schedules have their pains, either for the teams or for the customers Therefore: Reward teams for negotiating a schedule they can meet with financial bonuses [or at-risk compensation, or time off]. Keep two sets of schedules: One for the market and one for the team You cant wait until you have every last requirement to get started. Therefore: As soon as you have some confidence about project direction, start developing areas in which you have high confidence. Named Stable Bases It is important t...

You can count on me

You can count on me to have studied best practices You can count on me to have attended daily scrums You can count on me to have attended sprint planning meetings You can count on me to have attended sprint review meetings You can count on me to be prepared for all activities required to deliver a potentially shippable product You can count on me to present myself in a professional manner You can count on me to arrive at meetings at the designated time You can count on me to participate in the design meetings You can count on me to cooperate with the team goal during the sprint You can count on me to communicate You can count on me to be unafraid to question the actions of a fellow team member when I think an error might be made You can count on me to be able to accept information from a fellow member and admit that I was wrong You can count on me to have the courage to stick by my decision, even when a fellow member thinks I'm wrong You can count on me to be able to accept critici...

Scrum - Winner and Loser Attitudes

Surround yourself with winners! WINNER LOSER Wants the challenges to come his way Wants challenges to be dealt with by others Wants the sprint to be close and decide just in time what to re-negotiate. Wants the sprint goal to be a walk in the park Eager to tackle the difficult tasks Doesn't want the difficult problems Wants to be involved in team discussions, if needed Stays away from team discussions Wants to get it right Wants to avoid taking the blame Is not afraid to voice his opinion "It's their problem;" - not our problem Knows and understands best practices Doesn't know best practices Uses common sense No feel for the tasks Works hard Lazy Supports and defends team members "It's not my call." A team player . Only cares about himself Tries to help other team members "ME" not "we" attitude Makes other members better Has "rabbit ears" A leader A loner Works with product owner / Stakeholder Easily intimidated by prod...

Scrum Master vs Project Manager

Responsibility! That seems to be the key issue when moving to an agile mindset. In the beginning I was just confused and frustrated with the resistance. I did not understand how to handle it, nor could I explain why the difference was important and made sense. Now - quite some frustrating moments later - I'm ready to give explaining it a try. Traditionally the Project Manager is responsible for everything. The PM shall develop time plans, scope definitions, milestones, test plans, risk plans - well everything because it is the PM's responsibility! Does this actually mean that the PM spends all his time in the office creating these documents? No - or at least no good PM does. So what happens then? The PM - burdened with the responsibility - goes to the team and asks them to provide the information. Depending on the maturity of the team (and the PM) the PM then has some editorial tasks to compile the information. Then the PM states that the things are done by him because he is r...

The Project managers Iron Maid

Image
Back in the old days it was the Iron Maid that was a tortureres choice. Todays project managers may suffer the same from the Project Iron Triangle. The triangle is accurate in describing the variables that we have to work with i.e. Cost, Time and Scope. Traditionaly the Scope is fixed and then you work with cost and time. Everyone is trying to tie the project to a fixed time and cost too. In Agile the same triangle is present but with a twist. Its easier to say that I want something to this price at this time. Then let the product owner worry about the ROI of that something and let the team togheter with the product owner try to figure out what (scope) maximises RIO under the fixed time and cost.

Experience of Customer Impact as Defect Rating Value

Customer Impact as Defect Importance The usage of the two defect tracking fields severity and priority is widespread and dominant in defect management circles. But there is a number of problems with these attributes – they allow for ambiguity and they two fields do not focus on what is truly important – customer satisfaction. Let’s face it – without customer satisfaction we are out of business pretty fast. Thus to focus on what is important we have chosen to reflect what is important the defects impact on the customer. But you might say that not all of our systems have a customer! Well think twice – in the case where your systems don’t have an obvious end customer (player/operator/support/…) the component is a part of a producer-consumer chain. Your component allows for some other software module to provide functionality and there you are – you do have a customer! To reflect the impact on the customer we have chosen the following five states: • Catastrophic If this issue isn’t resol...

Fat & Fragile

Never thought that I'd say this but - internal debiting isn't just evil! We have had a situation where money really hasn't been one of the major issues when developing things. Well being that fat has been great - but now things have changed. To the better in some sense. Our product owners has when faced with priority issues always just blamed us for not recruiting fast enough. What I'm wondering is - have we focused on the right things from a business perspective or have we wasted valuable development time on chrome tires? This is where I think that internal debiting actually is a god and not a bad. Everything is useful and tasty when for free! But if you actually have to pay for it - does it carry the same appeal? - probably not! Forcing the stakeholder to provide the cash makes their part so painstakingly much clearer. They only have so much money and they need to get a product developed that will earn back the investment. This implies that they better have to do the...

Can Agile Maintenance Maintain Agility?

Just a few days back we fell of the bandwagon and it hurt badly. I'd say our most agile department was forced to close down and put the 10+ games they had into maintenance. Then almost all of the resources was reassigned into core business areas. But nothing bad without some tiny bit of good. In our case it was that we finally got both time and a quite substantial incitement to deal with how we maintain our products and how we run projects also. The product owners has not been accountable for ROI, or it simply hasn't been possible to calculate if a great ROI was better than a huge ROI. From a sound software development standpoint the business has been too successful. But the tide has turned. I finally got everyone accepting that we have to be more restrictive when starting projects. Previously a project has been a pool of resources following every whim of the product owner. But in a competitive world you change the software to earn money . The majority of such changes are eit...

Performance Reviews & Metrics

Any company that thinks of itself as a real company has a Human Resource department. Any HR department wants to measure it employees in order to establish salary ranges, incentive programs and so on. So far none has realized and far less understood that software development in general and in an Agile environment specifically is team work and not easilly mapped onto a sales guys metrics.

Plans are worthless, but planning is everything. – Eisenhower

What to do when there is no Product Owner present and you dont have a prioritized backlog? What about the project running wild? We need this and that and those bafoons at marketing that wants that Macromedia Installer are crazy? What if it breaks and installs our product wrongly - we think this is a top priority for the project to develop an in-house patcher/installer so that we dont have to buy that expensive MacroMedia thing. And the CRM in it that they want they surely dont know what they want and besides it cant be that important?

Tools for Fools

Aaargh - tools! Cant live with'em cant live without 'em. This truly is a topic that can get me going. It has so many dimensions but lets start with the folliest of follies - inhouse tool development. Brrr - I bet most has been directly involved or seen it either consume resources wildly or decay into uselessness. So what happens - some one has a great idea. Sure but then you do it as a skunk work and since its not core to the business you never ever get time and resources to make it even remotly done. So you end up with a piece of code that has expectations, since its been talked about for ages and the idea is great so the outcome has to even better. But it never gets there and eventually it will decay into worthlessness. The only good thing with this is that there hasnt been a serious consumption on resources to produce - well nothing :-) The other scenario is even worse in my opinion - you actually make a product/project out of the inhouse tool. But since it isnt core busin...