Friday, 5 March 2010

Hammered!

Fridays are normally interesting days, aren't they? No interventions or new actions should be scheduled for Fridays, to allow people enjoying a quiet weekend. But quite often Fridays come with a surprise. This morning surprise was this monitoring plot in the Ganglia PBS page. The CPU farm at PIC was being invaded by a growing red blob of very cpu inefficient jobs. The plot at the bottom pointed us to the originator: atlas pilot jobs.
The ATLAS Panda web page is quite cool, indeed, but not extremely useful for a profane to dig into it.
It took us quite some time to realise that the source of these extremely inefficient jobs was just at the end of the corridor: our ATLAS Tier2 colleagues submitting Hammercloud tests and checking that very low READ_AHEAD parameters for dCache remote access can be very inefficient. Next time we will ask them to keep the wave a big smaller.

Monday, 1 March 2010

LHC is back!

On February, the 7th, the CMS collaboration received the final positive referee report and publication acceptance on their very first Physics Results publication. The paper reports on first measurements of hadron production in proton-proton collisions occurred during the LHC commissioning December 2009 period. The successful operation and fast fata analysis impressed the editors and the entire collaboration was congratulated... and a party followed afterwards at CERN! ;)

This paper is under publication in JHEP and others will follow. CMS went into a major water-leak repair during the Winter shutdown, and now we are ready for more data. In fact, the LHC has restarted operations this weekend, and a few splash events have been already recorded by CMS.


After twenty years of design, tests, construction and commissioning, now is time for CMS collaborators to enjoy the long LHC run. LHC, we are prepared for the beams!

January availability report


We started 2010 with a number of issues affecting our two main Tier1 services: Computing and Storage. They were not that bad to make us failing the availability/reliability target (we still scored 98%) but sure there are lessons to learn.
The first issue affected ATLAS and it showed up on Jan 2nd in the evening, when the ATLASMCDISK token completely filled up: no free space! This is a disk-only token, so the experiment should manage it. ATLAS acknowledged it had had some issue with its data distribution during Christmas. Apparently they were sending to this disk-only token some data that should have gone to tape. Anyway, it was still quite impressive to see how ATLAS was storing 80 TB of data in just about 3 days. Quite busy Christmas days!
The second issue appeared on the 25th Jan and was more worrisome. The symptom was an overload of the dCache SRM service. After some investigation, the cause was traced to be the hammering of the PNFS carried out simultaneously by some MAGIC inefficient jobs plus also inefficient ATLAS bulk deletions. This issue puzzle our storage experts for 2 or 3 days. I hope we have now the monitoring in place that helps us next time we see something similar. One might try and patch the PNFS, but I believe that we can suffer from its non-scalability until we migrate to Chimera.
The last issue of the month affected the Computing Service and sadly had a quite usual cause: a badly configured WN acting as blackhole. This time it was apparently a corrupted /dev/null in this box (never quite understood how it appeared). We made our blackhole detection tools stronger after this incident, so that it will not happen again.

Thursday, 18 February 2010

PIC goes to IES Egara (outreach activity)


Last Tuesday, the 16th-February, Dr. Josep Flix went to IES Egara to give an overview of CERN, the LHC and the Grid to latest course high school students. It was not an easy task arriving there: it was raining, I was carrying around 100 CERN brochures with me, some other PIC brochures, the laptop, and all this... driving my bike! After getting lost on the town and asking several locals, I finally arrived to the high school. "Wet", but in time. The students were really surprised hearing what we do at CERN: science and technology. In the end, I was lucky enough not to electrocute myself during the talk (remember the rain and me "wet") and then students were able to place very interesting questions, indeed, well after the talk... Yes, the dark holes creation also was raised there, which seems a quite general and spread issue. From here, I want to congratulate Physics professor Juan Luis Rubio, to keep his students interested in Physics and with a very good knowledge of Particle Physics. After the talk, we spent also a good time in a nice restaurant on the town. At that time rain was gone...

Friday, 29 January 2010

IES Sabadell visits PIC (outreach activity)

Yesterday we had the visit of around 80 students and 5 professors from IES Sabadell to PIC installations. Their academic field based in Informatics ("Cicles formatius de grau mig/superior") made the tour to be exciting and full of questions. The visit, conducted by Dr. Josep Flix, started with two talks held in the IFAE Seminar Room (next to PIC). The first talk was entitled "The LHC and its 4 experiments: a data stream to understand the Big Bang" and was presented by Dra. Elisa Lanciotti, who is the LHCb contact at PIC. The students placed very interesting questions related to Physics and the techonology used on the LHC, during and after the talk. The level of curiosity was amazing! Maybe, in part related to the preparation sessions prior the visit the professors made and the comprenhesive Elisa's talk. Well after, Dr. Josep Flix presented "The use of Grid Computing by the LHC". He is currently the CMS contact at PIC and the CMS Facilities/Integration coordinator. The talk also raised questions from the attentive audience.


After the talks we made a visit to PIC installations, so they could see how a Computing Center is built and managed. In groups of 15 people we showed them first the real-time views of what's actually occurring on the Grid: the nice visualization of the WLCG grid activity on Google Earth, the ATLAS concurrent jobs running at all their Tiers, the CMS overall data transfer volumes, the LHCb job monitor display, and a few local monitoring plots, like the batch system and LAN/WAN usages.

Then, the visit to the Computing Area itself started: we showed them the different kind of disk pools we have installed, which covered the SUN X4500 (we opened one, so they could see how disks are installed and can be easily replaced) and the new powerful DDN system that offers 2 PBs of disk space; our computational power based on brand new HP Blade systems; plus the two tape robots we have at PIC (around 3 PBs of data stored) and which are the tapes available on the market and how we use them. The students were impressed as well on the WAN and LAN capabilities, the latest improved with the acquisition of two new 10 Gbps switches.


So far, the morning was extremely fruiful. From PIC we want to thank the Professors (Gregorio, Fernando, Lino, Alberto, Alexandra) for their dedication and motivation they offer to their students. They enjoyed the visit and want to repeat it with other students from the school in two months from now. We are happy to receive them again! ;)

Wednesday, 16 December 2009

Last day of LHC running this year

After so much celebration of first days of LHC running, it is time today to celebrate the last day of LHC running... this year. In few hours the LHC will be switched of and accelerated protons will go on holidays until next year.
I has been a very nice and long awaited time since last 23rd November the experiments started taking collision data. Today the LHC goes on holidays, but the WLCG does not. This piece of distributed infrastructure we have been building in the last six years should stay up and running 24x7 so that the precious data taken can be processed, re-processed, re-re-processed and so on. Somebody said that "data can be equated with money that has value only if it is used and circulated". So this is what we will be doing in the next weeks: giving value to the LHC data. This will not yet be haunting the Higgs, but less sexy minimum bias soft QCD events... but still, LHC physics after all.
At PIC Tier-1 we will carefully look to the services to ensure maximum availability and efficiency.
For the moment, what can we say about PIC's performance during "the month in which the LHC started" (aka November 2009)? We just received this Christmas gift from the official WLCG availability reports:
  • PIC availability and reliability for OPS VO = 100%
  • For ATLAS VO: 98% availability and 100% reliability (only ATLAS Tier-1 with max score)
  • For CMS VO: 100% availability and reliability (FZK also got max score for CMS)
  • For LHCb: 98% availability and 99% reliability (only CERN got 100% for LHCb)