Wednesday, 16 December 2009

Last day of LHC running this year

After so much celebration of first days of LHC running, it is time today to celebrate the last day of LHC running... this year. In few hours the LHC will be switched of and accelerated protons will go on holidays until next year.
I has been a very nice and long awaited time since last 23rd November the experiments started taking collision data. Today the LHC goes on holidays, but the WLCG does not. This piece of distributed infrastructure we have been building in the last six years should stay up and running 24x7 so that the precious data taken can be processed, re-processed, re-re-processed and so on. Somebody said that "data can be equated with money that has value only if it is used and circulated". So this is what we will be doing in the next weeks: giving value to the LHC data. This will not yet be haunting the Higgs, but less sexy minimum bias soft QCD events... but still, LHC physics after all.
At PIC Tier-1 we will carefully look to the services to ensure maximum availability and efficiency.
For the moment, what can we say about PIC's performance during "the month in which the LHC started" (aka November 2009)? We just received this Christmas gift from the official WLCG availability reports:
  • PIC availability and reliability for OPS VO = 100%
  • For ATLAS VO: 98% availability and 100% reliability (only ATLAS Tier-1 with max score)
  • For CMS VO: 100% availability and reliability (FZK also got max score for CMS)
  • For LHCb: 98% availability and 99% reliability (only CERN got 100% for LHCb)

Sunday, 22 November 2009

Outreach in the school

Last week it was the "week of science" in Spain. This happens every year around mid November and consists of one week where plenty of activities oriented to explain science to the people are scheduled. In Catalonia, one of the organised activities are talks of scientists in the schools. Last wednesday there were 100 simultaneous talks carried out in different schools all around Catalonia. I visited a secondary school in Badalona where I had a great time talking about the LHC and the origin of the Universe to around 70 students. I see now that they even posted an etry in the blog of the school!
Nice to see that Catalan schools are in the blogosphere... and that they had a nice time listening to my LHC stories.

Real data flowing through PIC


ATLAS has provided a nice monitoring page where we can follow the progress of data distribution in these so exciting moments of first circulating beam in the LHC. This is not collisions yet, but real data indeed. After so many years of simulations, we are happy to see the first Megabytes of real stuff. In the picture, I have just captured the current status of the datasets distribution to Tier-1s and from there to the associated Tier-2s. The overall picture looks pretty green, which is good news. PIC received the subscribed data with no problems and promptly redistributed it to the Tier-2s. It looks the data movement went mostly smooth. Let's keep an eye on this. We will see the rates growing in the next days.

Circulating beam in the LHC (take two)


So, there we go. Last friday 20th November beams circulated again inside the LHC, after one long year of reparations. Everyone is happy and bottles of champain (or cava) are being opened in the control rooms. In the picture you can see, besides the party atmosphere at the LHC control room, the first event displays from ATLAS, CMS and LHCb . The hundreds of tracks coming from the collimators where beams are splashed can be clearly seen in all of them. We are watching the first LHC data.
Commencing countdown, engines on...

Tuesday, 10 November 2009

LHC beam approaching CMS!

Last Saturday evening, the 7th of November 2009, at around 8 p.m., after passing through the LHCb detector, for the first time since last year's incident, protons arrived at the doorstep of the CMS experiment, thus completing half the journey around the LHC's circumference.

Low energy protons from the LHC were dumped in a collimator just upstream of the CMS cavern. The calorimeters and the muon chambers of the experiment saw the tracks left by particles coming from the dumping point (a so-called 'splash event', see images). During the rest of the weekend, bunches of protons were also sent in the clockwise direction passing through the ALICE detector and were dumped at point 3.

All detectors saw 'splash' events on their monitoring pages. Castor and the Preshower detectors saw particles for the first time! Some beautiful pictures from the events seen:

Monday, 19 October 2009

Data loss


These days we are hearing more often about data loss events at the WLCG sites. Today it was the NL-T1 site that reported some data loss in the daily Operations meeting. Apparently, the tape drive loaded the tape and, instead of reading it, it just destroyed it. A similar event happened to us at PIC at the end of September, when we lost a tape containing 214 files from CMS. Nothing could be done with that piece of hardware... not even rewinding it! Luckily for us, all of those files were replicated in some other Tier-1 or at CERN, so we could fix the problem quite straightaway.
We were used to think in tapes as a safe media for data... but these episodes of tape destruction show that this is not always the case. A bit scary.
Anyhow, even ig it does not solve anything but it is nice to see that WLCG sites we are not the only ones losing data. Even Microsoft loses some data eventually!

Friday, 25 September 2009

August availability: on target, but...

Last August, as may be one could expect after a busy July with the CPU occupied at almost 100%, many of the LHC experiment production managers and computing guys went on (deserved) holidays. Accounting data is being collected these days, and the results for PIC show that only about 40% of the installed CPU was used. Curiously, the experiment share of this workload was highly non-nominal: LHCb, which represents only about 10% of PIC pledges, was the one consuming more: up to 60% of the delivered CPU cycles were for them. ATLAS got essentially all the remaining 30%, while CMS did essentially zero. So, for both ATLAS and CMS August was the month with the lowest CPU consumed in the year, while for LHCb was a record-breaking month.
PIC Tier-1 availability during August was just right on top of the target: 97%. About 1% of the unavailability was due to the usual monthly Scheduled Downtime which took place on the 25th August. Most of the remaining 2% unavailability we spent it also on that same day, which suggests that there is still room for improvement on SD coordination. One of the A-critical services, the site-bdii, was off during almost 4h after its scheduled intervention... and no one of us noticed! There was also an issue with the Computing Service (B-critical) which had its queues closed for 2h longer than planned. We should now then feed this experience back into our operation system and make sure the relevant procedures are improved.
Besides that, the 19th of August instabilities appeared in the OPN link which were affecting the SRM service, specially for outgoing transfers. The problem disappeared in about 24h, but we never knew what had really happened. The Spanish NREN did not answer to our query for information. First we thought this was an August-effect, but later we realised the problem was that our e-mail contact for operational issues in the network was wrong. We have corrected this and the e-mail we have now should even trigger a ticket opening automatically.
The good news for August were that, despite being one of the hottest in several years, the cooling system of the PIC machine room coped perfectly with it. Seems that the new maintenance team did a good job in preparing the system for the summer campaign.