Tuesday, 10 November 2009

LHC beam approaching CMS!

Last Saturday evening, the 7th of November 2009, at around 8 p.m., after passing through the LHCb detector, for the first time since last year's incident, protons arrived at the doorstep of the CMS experiment, thus completing half the journey around the LHC's circumference.

Low energy protons from the LHC were dumped in a collimator just upstream of the CMS cavern. The calorimeters and the muon chambers of the experiment saw the tracks left by particles coming from the dumping point (a so-called 'splash event', see images). During the rest of the weekend, bunches of protons were also sent in the clockwise direction passing through the ALICE detector and were dumped at point 3.

All detectors saw 'splash' events on their monitoring pages. Castor and the Preshower detectors saw particles for the first time! Some beautiful pictures from the events seen:

Monday, 19 October 2009

Data loss


These days we are hearing more often about data loss events at the WLCG sites. Today it was the NL-T1 site that reported some data loss in the daily Operations meeting. Apparently, the tape drive loaded the tape and, instead of reading it, it just destroyed it. A similar event happened to us at PIC at the end of September, when we lost a tape containing 214 files from CMS. Nothing could be done with that piece of hardware... not even rewinding it! Luckily for us, all of those files were replicated in some other Tier-1 or at CERN, so we could fix the problem quite straightaway.
We were used to think in tapes as a safe media for data... but these episodes of tape destruction show that this is not always the case. A bit scary.
Anyhow, even ig it does not solve anything but it is nice to see that WLCG sites we are not the only ones losing data. Even Microsoft loses some data eventually!

Friday, 25 September 2009

August availability: on target, but...

Last August, as may be one could expect after a busy July with the CPU occupied at almost 100%, many of the LHC experiment production managers and computing guys went on (deserved) holidays. Accounting data is being collected these days, and the results for PIC show that only about 40% of the installed CPU was used. Curiously, the experiment share of this workload was highly non-nominal: LHCb, which represents only about 10% of PIC pledges, was the one consuming more: up to 60% of the delivered CPU cycles were for them. ATLAS got essentially all the remaining 30%, while CMS did essentially zero. So, for both ATLAS and CMS August was the month with the lowest CPU consumed in the year, while for LHCb was a record-breaking month.
PIC Tier-1 availability during August was just right on top of the target: 97%. About 1% of the unavailability was due to the usual monthly Scheduled Downtime which took place on the 25th August. Most of the remaining 2% unavailability we spent it also on that same day, which suggests that there is still room for improvement on SD coordination. One of the A-critical services, the site-bdii, was off during almost 4h after its scheduled intervention... and no one of us noticed! There was also an issue with the Computing Service (B-critical) which had its queues closed for 2h longer than planned. We should now then feed this experience back into our operation system and make sure the relevant procedures are improved.
Besides that, the 19th of August instabilities appeared in the OPN link which were affecting the SRM service, specially for outgoing transfers. The problem disappeared in about 24h, but we never knew what had really happened. The Spanish NREN did not answer to our query for information. First we thought this was an August-effect, but later we realised the problem was that our e-mail contact for operational issues in the network was wrong. We have corrected this and the e-mail we have now should even trigger a ticket opening automatically.
The good news for August were that, despite being one of the hottest in several years, the cooling system of the PIC machine room coped perfectly with it. Seems that the new maintenance team did a good job in preparing the system for the summer campaign.

Friday, 18 September 2009

One Petabyte knocking at the door

This is just a micro-post, to show you who I found two days ago knocking at PIC's door just when I was leaving the office.
Yes! the long-awaited Petabyte that will push our capacity up to the 2009 MoU pledges was there.
It will be a tricky path the one we have to walk from these nice wooden boxes to the WLCG SRM service (the first one seems to be the big disks are reluctant to get to Barcelona).
Busy weeks ahead... but the goal is clear: fill up those disks with LHC data.

Friday, 28 August 2009

Believe us, PIC is Ok


Scary, isn't it? PIC availability has been bully red in the last 24h but, sadly, there was not much we could do.
We have always said, and still strongly believe it, that SAM tests are a very good thing. Actually, I firmly believe that have been one of the key ingredients for the WLCG success. Success here meaning the evolution from "the Grid does not work" situation with 60% job success rate we had few years ago, to the rutine >97% availabilities we are used to see these days. But yes, not even SAM tests are perfect. There has always been a dark corner inside them: the so much questioned "SE test inside the CE", or lcg-rm test. And inside this controversial test there is another smaller corner which is still a bit darker: the file replication to CERN test. This was the one that started flickering on Tuesday at PIC and it is consistently failing since more than 24h. This test tries to copy a file sitting at PIC into a very concrete DPM server at CERN. This very precise connection was timing out for us while any other transfer to any other site, even to any other CERN storage server was working. This was strange enough so that we asked for help to our CERN colleagues. Today, they came with the good news: problem found, a problematic router.
Got a really puzzling error? Bet on the network...

Friday, 21 August 2009

July: good availability comes with CPU delivery record


Most of the people is these days either just back from holidays (me) or still out (lucky them) or neither of those (also lucky, since they will leave later...). Anyhow, life at the Tier-1 is 24x7 since we all know so, it doesn't matter if we are in the middle of August and it is near 40 degrees celsius out there, it is time to report about service performance in the past month.
We just got the WLCG reports for July and the results for PIC are pretty positive: 99% availability and reliability. This little 1% that gets us away of our beloved 100% happened on the 29th July, around noon. During four hours all of the Computing Elements at PIC were failing the SAM Job Submission tests, so definetely the Tier-1 service was affected during that period. The source of the problem was found to be a pretty mysterious one: the switch connecting the servers hosting Virtual Machines to the PIC LAN (VMs are always a bit misterious, aren't they?). Actually, the source of the problem was not found, but just disappeared when that switch was replaced by a new one (different brand, no names here to avoid anti-propaganda :)
Regarding the availability of PIC as seen from the experiment specific monitoring, we got also very good results for ATLAS and LHCb, close to 100%. However, the result for CMS was not that good: a mere 90%. This funny assymetry was due to a bug we introduced by mistake in the Torque ACL queues configuration which actually blocked CMS submissions to our short queue for 3 days (6-9 July). Somebody could ask why we did not notice this in 3 days... We should put priority in deploying the WLCG-Nagios in production. It will for sure help reducing these unavailable times.
So, besides this unfortunate CMS-blocking bug, we can say July was a pretty good month for the Tier-1 in terms of availability. Thanks to this, and also to the job submission hyperactivity seen from ATLAS and LHCb, we delivered a record amount of CPU cycles during July: around 70.000 ksi2k·days, which is very close to keeping 100% of our resources busy for the whole month.
Now things look quiet, being many people still away. PIC is up and running, cooling is ok (even the heat wave out there)... but watch out for the VM-networking ghosts, they could come at any time.

Thursday, 18 June 2009

Welcome Valparaiso !

The Universidad Técnica Federico Santa Maria (UTFSM) of Valparaiso (Chile) has joined ATLAS Computing and has been associated to PIC as its Tier-1 center. They are now a small center (Tier-3) but surely will grow soon. Firsts transfers using the full chain of the ATLAS Data Management System has been successfully tested. We welcome UTFSM to PIC and to the distributed computing world !