Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Thursday, September 6, 2018

How to interpret the literature: A new series of posts

The Salty Statistician will be a recurring feature of this blog wherein we ask statisticians in medicine to break down articles from the surgery literature and assess whether the reported conclusions are supported by the data. Let’s look at this study:

Groh MA et al. Is Surgical Intervention the Optimal Therapy for the Treatment of Aortic Valve Stenosis for Patients With Intermediate Society of Thoracic Surgeons Risk Score? Annals of Thoracic Surgery.

The authors attempted to address the question of whether aortic stenosis patients deemed “intermediate risk” [IR] for surgical aortic valve replacement [AVR] are best treated with open surgery or transcatheter AVR. The authors looked at 1,144 patients who received surgical AVR from 2008-2014 at a single center focusing on the 620 “intermediate risk” patients. At the end of the follow-up period, 72 had died.

Unfortunately, major methodological issues undermine the paper’s conclusions. Fortunately, this provides an excellent teaching opportunity.

First, the authors inappropriately used logistic regression to analyze independent predictors of mortality. Logistic regression treats the outcome as a simple “Yes” or “No” variable, while ignoring the time-at-risk. This study included patients treated over a six-year period (2008-2014) who therefore have substantial differences in the amount of time at risk. Consider the following hypothetical patients.

Patient A treated in 2008 and died in 2014 surviving six years after surgery. The logistic regression model simply counts patient A as “dead.”

Patient B treated in 2014 and alive in 2017 but dies in 2018, after the data were analyzed and the paper published. He survived four years after surgery and in the logistic regression model, counts as “alive” since data were analyzed in 2017.

Patient A lived for six years after surgery, but counts as “worse” in the analysis than Patient B who only lived for four years because of the time at which the data were “frozen” and analyzed. Of course, this is unavoidable in long-term outcomes studies, but one must choose an appropriate statistical method that accounts for time-at-risk.

Cox proportional-hazards models are more appropriate for a long-term survival outcome than logistic regression. When building a Cox model, one specifies both the current status (i.e., alive/dead) as well as an amount of follow-up time. For example, Patient A is “dead” with six years of follow-up; Patient B is “alive” but with only three years of follow-up. This provides a proper assessment of how strongly the independent variables are associated with risk of mortality while accounting for the unequal follow-up time.

Second, the authors state their data supports the conclusion that “SAVR is the optimal therapy for most of the patients” in the IR group in comparison to TAVR. However, their paper lacks any data on outcomes in IR patients who were treated with TAVR. Why the authors believe presenting data from a series of SAVR patients is sufficient to claim that SAVR is the “optimal therapy” absent any comparison data on patients treated with TAVR is unclear. Randomized controlled trials have more appropriately compared SAVR and TAVR in the IR population. Link here and here.

Which patients should receive surgical AVR versus transcatheter AVR is a good question, but to answer it, the paper used an incorrect approach.

Final Rating (1-5 Scalpels): 1 Scalpel - significant methodological issues

This issue of the Salty Statistician was written by Andrew Althouse (@ADAlthousePhD), currently an Assistant Professor of Medicine at the University of Pittsburgh as well as Statistical Editor of Circulation: Cardiovascular Interventions.

We intend this series to focus on work that is perceived to have a high impact on clinical practice, so we welcome reader suggestions. If you have a paper that you would like to see reviewed as part of the Salty Statistician series, please tweet @Skepticscalpel or @ADAlthousePhD or email SkepticalScalpel@Hotmail.com. We cannot promise that all submissions will be reviewed in this space, but we will do our best.


Tuesday, June 26, 2018

We need less research

“We need less research, better research, and research done for the right reasons. Abandoning using the number of publications as a measure of ability would be a start.” Although I have expressed similar sentiments in blog posts [here and here], I didn’t say it. It was written by Douglas Altman, a well-known statistician and researcher who died in June.

Altman made that statement in a 1994 BMJ article entitled “The scandal of poor medical research.” Here we are, 24 years later, and nothing has changed. In fact, thanks to the rise of predatory journals, things are much worse.

Altman lamented research containing flaws such as “the use of inappropriate designs, unrepresentative samples, small samples, incorrect methods of analysis, and faulty interpretation” and felt many poor studies were the result of pressure on researchers to publish.

Monday, January 15, 2018

Facial exercises to make you look younger? I don't think so.

One would think a study covered by the New York Times would be both scientifically valid and important. Apparently, that is not always the case.

Under the headline “Facial Exercises May Make You Look 3 Years Younger” is a story about a research letter published in JAMA Dermatology. The Times article concludes with a quote from the lead author, “But for now, it is reasonable to consider contorting and pinching up your face if you wish to try to look younger.”

Is it reasonable? Let’s see if this 1½ page research letter proved its point.

Friday, October 21, 2016

The moon and hospital admissions

A few days ago we had a full moon. A lengthy discussion about the effect of a full moon on hospital admissions took place on Twitter.

Many papers say admissions increase and odd things happen, and many others have found there is no relationship between the phases of the moon and anything that goes on in hospitals.

Someone sent me a link to a paper that a lot of devotees of astrology like to quote. It's called "The influence of the full moon on the number of admissions related to gastrointestinal bleeding," and it appeared in the International Journal of Nursing Practice in 2004.


Tuesday, June 7, 2016

Changing pre-med requirements and med school curricula

Ezekiel Emanuel, the University of Pennsylvania physician and ethicist, has written an opinion piece suggesting many changes in both pre-medical education and the medical school curriculum.

He would do away with many of our hallowed medical school prerequisites such as calculus, physics, and organic chemistry, feeling that those subjects are simply used to "weed out" certain students. I confess I once believed that such subjects were worthwhile. However, Emanuel makes a convincing argument that rigorous college courses in more relevant disciplines such as statistics, genetics, ethics, and psychology with a special focus on human behavior would suffice.

Regarding medical school, Emanuel points out he was taught the Krebs cycle on four different occasions in college and medical school and never used it once in practice or research. I have made a similar observation in a previous blog post.

He considers pathology, cytology, and pharmacology to be largely irrelevant to medical practice but concedes that some may disagree.

Tuesday, March 8, 2016

An intraoperative leak test should not be done; or should it?

Here is an abstract recently published ahead of print in the American Journal of Surgery. Please read it because a one-question test follows.

Introduction: Staple line leak after sleeve gastrectomy (SG) is a rare but dreaded complication with a reported incidence of 0-8%. Many surgeons routinely test the staple line with an intraoperative leak test, but there is little evidence to validate this practice. In fact, there is a theoretical concern that the leak test may weaken the staple line and increase the risk of a postop leak.

Methods: Retrospective review of all SG performed over a 7-year period. Cases were grouped by whether an intraoperative leak test (IOLT) was performed, and compared for the incidence of postop staple line leaks. The ability of the IOLT for identifying a staple line defect and for predicting a postoperative leak was analyzed.

Results: 542 SG were performed between 2007-2014. 13 patients (2.4%) developed a postop staple line leak. The majority of patients (N=494, 91%) received an IOLT, including all 13 patients (100%) who developed a subsequent clinical leak. There were no (0%) positive IOLTs and no additional interventions were performed based on the IOLT. The IOLT sensitivity and positive predictive value were both 0%. There was a trend, although not significant, to increased leak rates when a routine IOLT was performed versus no routine IOLT (2.6% vs. 0%, p=0.6).

Conclusions: The performance of routine IOLT after sleeve gastrectomy provided no actionable information, and was negative in all patients who developed a postoperative leak. The routine use of an IOLT did not reduce the incidence of postop leak, and in fact was associated with a higher leak rate after SG.


Do you agree with the authors that the routine use of the IOLT was associated with a higher leak rate after sleeve gastrectomy?

Thursday, January 14, 2016

Can patients shower immediately after surgery?

Here’s what a recent paper published ahead of print in Annals of Surgery says:

Between May 2013 and March 2014, 222 patients were randomized to the group allowed remove their dressings and shower at 48 hours and 222 to the group permitted to shower only after the original dressing and the sutures were removed in clinic. There were 4 (1.8%) superficial surgical site infections in the early shower group and 6 (2.7%) in the late shower group, an insignificant difference with p = 0.751.

The authors concluded that clean and clean-contaminated wounds can be safely showered 48 hours after surgery, and early postoperative showering may increase patient satisfaction.

I have always been an advocate of early showering after surgery. Wounds properly closed will be bridged by epithelium within 48 hours. Tap water is relatively sterile or we couldn't drink it. Many studies have shown that even irrigating open wounds with tap water instead of sterile saline does not lead to more infections. [Links here and here.]

Much as I would like to believe the Annals study, I can’t because it is probably underpowered to show a difference between the two groups.

Here is a nice definition of statistical power from a website called effectsizeFAQ.com:

“In plain English, statistical power is the likelihood that a study will detect an effect when there is an effect there to be detected. If statistical power is high, the probability of making a Type II error, or concluding there is no effect when, in fact, there is one, goes down.”

To their credit, the authors did try to estimate the sample sizes they would need by doing a power calculation. They knew that the wound infection rate for the cases they intended to enroll was about 1%. The problem is they estimated that showering at 48 hours would result in a wound infection rate of 5%. That seems very high to me for the types of cases included in their investigation—thyroid, lung, inguinal hernia and skin tumors.

If they had hypothesized that early showering would merely triple the rate of wound infections from 1% to 3%, they would have needed at least 1536 patients in each arm of the study. Then if there was no difference, one could conclude that early showering truly does not cause more wound infections.

Even if the known incidence of wound infection was much larger, say 5%, and the rate of infection with showering was presumed to be doubled (10%), to have enough power a study would need 434 patients in each arm.

Many websites provide calculators for determining the appropriate sample sizes to detect with a reasonable degree of certainty whether one intervention is better than another. Anyone thinking about doing a prospective randomized trial should realistically estimate the expected difference and calculate the power.

Whenever you read a negative study, the first question to ask is, “Was the study adequately powered to avoid a type II error?”

Wednesday, September 2, 2015

Variation is not causation

I made a rookie mistake in statistics of the “correlation is causation” genre by confusing variation for causation in the recent JAMA Surgery paper referred to in my last post. I contacted Dr. Timothy M. Pawlik, the lead author of the Johns Hopkins study, who said the following:

"The model is explaining and attributing variation in readmission and not attributing readmission itself to the different domains. The model suggested that only 2.8% of the variation in readmissions was attributable to surgeons. This is different than saying that only 2.8% were the 'fault' of surgeons. A more accurate interpretation would be that only 2.8% of the variation seen in readmissions was attributable to provider level factors. The majority of the variation in readmission was due to patient factors."

He added that some of the 82.8% variation in readmissions attributable (note: attributable doesn’t mean it’s the patient’s fault) to the patient could be modified by better medically managing patients' comorbidities or not operating on some of these patients.

That readmissions can be explained by a single domain or a single person is simplistic. Dr. Pawlik's clarification confirms my original concern that attributing differences in patient outcomes solely to differences in technical quality of surgeons is probably inaccurate, statistically speaking.

Variation is not causation but variation is still a call to action. Regardless of who is to blame for unfavorable outcomes, surgery is a team sport. The incision is just as important as the community care. In this regard, I am certain that ProPublica and I are on the same side. Let’s work together so that we see the whole story behind the numbers.



Friday, July 24, 2015

The Surgeon Scorecard: My analysis

I've got nothing against ProPublica. If a valid way to rate surgeons is ever discovered, I would support it completely. However, ProPublica's Surgeon Scorecard is not the answer.

I keep hearing its defenders say, "Some data is better than no data at all." I disagree strongly with that. To me, bad data is worse than no data at all. People with much more statistical sophistication than I have pointed out the flaws in the scorecard.

Digression: Having written many posts about statistics, I can tell you that the mere mention of the word drives readers away about as fast as if you were to yell "Fire" in a crowded theater.

I want to focus on a different area. The scorecard has created a lot of chatter on Twitter, and just about everyone I know has blogged about it.

This reminds me of a couple of posts I wrote back in 2011. [Links here and here.] I pointed out that Twitter might not be as important as those of us who use it think it is.

While we were busy arguing about the merits of the scorecard on Twitter, I'm not so sure what the general public was doing.

For example, ProPublica says the Surgeon Scorecard has had over 1 million visitors since its launch. That sounds like a lot until you consider that the current population of the United States is estimated at 321 million. So 1 million people would be 0.3%. We do not know how many of those 1 million were unique visitors. It could be that many of them were doctors looking for their own statistics and bloggers looking for ideas.

That the public may not care was reinforced by a rather tepid response to the ProPublica AMA (Ask Me Anything) on Reddit today.

By 1:00 PM EDT, which was two hours into the AMA, there were 80 comments, 31 of which were by ProPublica staff or the spine surgeon who had consulted on the scorecard's methods.

Just to give you some perspective, an AMA last year by a guy with two penises drew 17,134 comments.

Because the demographic is skewed toward younger people, perhaps Reddit may not have been the right venue. Although Reddit boasts 169 million unique visitors per month, the most recent figures show that 33% of the Reddit users are mostly men between 18 and 49 years old. Those under 18 are not counted but represent "a substantial percentage of Reddit users."

My two favorite questions asked of ProPublica were "How can I tell if my doctor is capable of making an error?" and "Do you fix the leg which is broken completely?" [Did the question refer to a leg that was completely broken, or did it mean should the leg be completely fixed?]

What have we learned here? It's hard to say.

If you want to read a measured critique of the scorecard, go to Dr. John Mandrola's piece on Medscape.

Thursday, February 19, 2015

Don't trust the abstract; read the whole paper


Above are the results and conclusion from an abstract of a paper called "Single-Incision Laparoscopic Cholecystectomy: Will It Succeed as the Future Leading Technique for Gallbladder Removal?" It appears online in the journal Surgical Laparoscopy, Endoscopy & Percutaneous Techniques.

It is another great example of why you need to read the entire paper and not just the abstract.

The methods section says that the study involved 875 patients with prospectively collected data. Don’t be fooled by prospectively collected data. This research is retrospective.

Here are some issues.

Saturday, October 11, 2014

Is student test performance impaired by distracting electronic devices?

After listening to a lecture, third-year students at the Harvard School of Dental Medicine were surveyed about distractions by electronic devices and given a 12-question quiz. Although 65% of the students admitted to having been distracted by emails, Facebook, and/or texting during the lecture, distracted students had an average score of 9.85 correct compared to 10.444 students who said they weren't distracted. The difference was not significant, p = 0.652.

In their conclusion they authors said, "Those who were distracted during the lecture performed similarly in the post-lecture test to the non-distracted group."

The full text of the paper is available online. As an exercise, you may want to take a look at the paper and critique it yourself before reading my review. It will only take you a few minutes.

As you consider any research paper, you should ask yourself a number of questions such as are the journal and authors credible, were the methods appropriate, were there enough subjects, were the conclusions supported by the data, and do I believe the study?

Monday, September 8, 2014

Chance can turn a surgeon into a killer

Risk-adjusted 30- to 90-day outcome data for selected types of operations done by specific surgeons and hospitals are now being publicly posted online by England's National Health Service.

According to the site, "Any hospital or consultant [attending surgeon in the UK] identified as an outlier will be investigated and action taken to improve data quality and/or patient care."

After cardiac surgery outcomes data were made public in New York, some interesting unexpected consequences were noted.

Surgeons and hospitals resorted to "gaming the system" by declining to operate on patients who were high-risk and tinkering with patient charts to make those they did operate on seem sicker. This can be done by scouring the charts for all co-morbidities and making sure none are overlooked when they are coded. An article from New York Magazine explains it in more detail.

Interpreting outcomes data can be tricky.

In a post three years ago about a report that nine Maryland hospitals had higher-than-average complication rates, I pointed out that whenever you have averages, some hospitals are going to be worse than average unless all hospitals perform exactly the same way or, like medical students, are all above average.

A much more sophisticated way of looking at this subject appeared in a fascinating 2010 BBC News piece by Michael Blastland, who is the Nate Silver of England [or maybe Nate Silver is the Michael Blastland of the US], called "Can chance make you a killer?"

Blastland set up a statistical chance calculator for a hypothetical set of 100 hospitals or 100 surgeons performing 100 operations each. The model assumes that every patient has the same chance of dying and that every surgeon is equally competent. The standard is that a mortality rate 60% worse than the norm set by the government for any hospital or surgeon is not acceptable.

You are assigned one hospital. Using a slider, you may choose an operative mortality rate anywhere from 1% to 15%. After you do this a number of times and recalculate for each mortality rate, you will notice that the number of unacceptably performing hospitals or surgeons changes randomly for each percent mortality and your hospital may appear in the underperforming group strictly by chance alone.

The whole concept is explained in more detail on the site. I encourage you to try it for yourself. The link is here.

So it may be difficult for the NHS to separate the true outliers from the unlucky surgeons who happened to fall outside the established norms.

What do you think about this?

Friday, June 27, 2014

Check the Y-axis when reading a chart


Here is an interesting way to make statistics more persuasive without really lying.

Take a look at this chart. It was labeled "Hospital readmissions sharply declined."


Now look at this one. The same data are charted, but the decline does not look nearly as sharp.


A common trick is to abbreviate the Y-axis of a chart. Proponents of this will tell you it makes the chart more compact and easier to read. However, the downside is that a small change is made to appear much larger.

Even though the change was said to be statistically significant in this instance, the casual reader would certainly be impressed much more by the sharp decline depicted in the first chart.

I believe the first chart was produced by the Centers for Medicare and Medicaid Services (CMS) to show that its policy on readmissions was working. I was unable to find the original source for it.

This also works with other graphics as shown in this pair of bar charts from the Visualizing Data blog.


I encourage you to look for this type of manipulation when you read research papers. pharmaceutical ads, or any other depiction of data.




Friday, May 23, 2014

Cows or Sharks? Which are more likely to kill you?

Give me a minute, and I'll get to the cows and sharks.

You would be surprised at how few doctors are familiar with even the most basic statistics. Medical journal articles often have statistical errors which are missed by manuscript peer reviewers and readers alike.

Most medical students have taken a course in statistics, but it is usually taught in the first or second year of school. By the time they start residency training when they could really use the information, they have forgotten most of it. Statistics should be taught during the clinical years of medical school and reinforced throughout residency training.

Hospital administrators are even more clueless than physicians. I have blogged before (here and here) about the irrational responses of administrators to miniscule changes in poorly constructed surveys of patient satisfaction. When scores go down by insignificant percentages, all hell breaks loose with task forces, ad hoc committees and browbeating of staff.

Here’s a fun exercise involving statistics. It's OK. No formulas will be discussed.

Which animal kills more people per year in the United States, cows or great white sharks?

Although not long ago a German tourist was killed by a shark in Hawaiian waters, the answer is overwhelmingly "cows."

How can this be? You rarely hear about a cow killing a human but it happens about 20 times every year. Between 2003 and 2008, 108 people died from injuries caused by cattle across the United States, according to the Centers for Disease Control and Prevention. That's 27 times the whopping 4 people killed in shark attacks in the United States during the same time period, according to the International Shark Attack File.

Guess how many cows there are in the US.

According to the Drovers Cattle Network, there were 96.5 million head of cattle here as of mid-2013. The cattle population dwarfs the number of great white sharks. The New Ecologist estimates that the number of great whites in the entire world is about 3500.

The Guardian recently reported that there have been 1,085 recorded shark attacks in the US since the year 1670 for an average of only 3.5 shark attacks each year for the last 342 years.

Although not as dramatic or as newsworthy as a shark attack, it is far more likely that a person will be killed by a cow than a shark.

So keep your statistical radar turned on. Be skeptical.

And if you see an udder in the water, get to shore as fast as you can.


Wednesday, March 19, 2014

A study says you can trust online physician ratings

This abstract comes from the Social Science Research Network:

Despite heated debate about the pros and cons of online physician ratings, very little systematic work examines the correlation between physicians’ online ratings and their actual medical performance. Using patients’ ratings of physicians at RateMDs website and the Florida Hospital Discharge data, we investigate whether online ratings reflect physicians’ medical skill by means of a two-stage model that takes into account patients’ ratings-based selection of cardiac surgeons. Estimation results suggest that five-star surgeons perform significantly better and are more likely to be selected by sicker patients than lower-rated surgeons. Our findings suggest that we can trust online physician reviews, at least of cardiac surgeons.

You won't be surprised to learn that I don't believe it. As is my custom, I decided to read the entire paper the full text of which can be found here. At 37 pages, the raw manuscript is rather lengthy. As a public service, I waded through it.

The authors, non-MD faculty from the William E. Simon Graduate School of Business Administration at the University of Rochester, in New York, combed the ratings for Florida cardiac surgeons on the website RateMDs.com and classified surgeons into three categories—five-star surgeons, non-five-star surgeons, and those with no ratings at all.

They looked at 799 quarterly opportunities for ratings over a 9-year period and found that 21% of surgeons had an average of 1.9 online ratings. The 79% of surgeons who did not have an online rating performed 79% of the total surgeries in 2012, the year that the authors analyzed for patient results.

The five-star surgeons had a mean of 1.8 reviews each, and only 10% had more than 2 reviews.

The average mortality rate for coronary artery bypass grafting (CABG) among the Florida cardiac surgeons was 1.8% in 2012. The five-star surgeons with multiple reviews had the highest mortality rates at 3.3%.

I could find no evidence that patient mortality rates were adjusted for risk. But a lot of statistical manipulations took place. It's all explained by this simple equation—one of many.

 The authors say, "For a representative patient who is severely ill, being treated by a five-star surgeon can reduce the in-hospital mortality by 55% compared with being treated by a non-five-star surgeon. [I have no idea how they determined that figure.] Moreover, the negative and significant coefficient of no-ratings suggests that patients treated by surgeons without ratings also have a lower mortality rate than those treated by non-five-star surgeons, all else being equal." Huh?

And this, "Patients with private insurance are less likely to select the surgeons without ratings than patients with Medicare. We suspect that patients with private insurance have to use search engines to figure out whether a surgeon is within the network that an insurance plan covers, while government patients enjoy a large physician network." I question that assumption. My experience is that patients with Medicare sometimes have problems finding anyone to care for them, let alone the best surgeons.

It turns out that half of the five-star surgeons had only one review. In one iteration of the study model, five-star surgeons with multiple reviews had higher mortality rates than those with only one review, but then they also say, "One surprising finding is that five-star surgeons with a single review show no statistical difference in performance from those with multiple reviews."

Are you as confused as I?

The paper makes no mention of the possibility that some of the online ratings could be fake. Recent articles [here and here] suggest that one-fifth to one-third of such reviews are phony.

You can manipulate the statistics all you want, but you won't convince me that one or two or even 20 online ratings are valid or useful in choosing a surgeon.

Tuesday, February 25, 2014

"Medical errors kill hundreds of thousands each year in the US"


How about that headline?

It appeared on RT.com, "the first Russian 24/7 English-language news channel which brings the Russian view on global news."

The story, which originally ran in November of 2013, was resurrected again on Twitter yesterday. It's subject was a paper that claimed as many as 440,000 patients die from medical errors in the United States every year.

Back in September, I criticized the study because it assumed that every death was both preventable and caused by a medical error. Neither assumption is correct. It also extrapolated the doomsday figures from only four other papers describing just 38 deaths.

In that post I said, "Inflating the incidence of these problems does nothing but further erode the already shaky confidence of the public in the medical profession. And creating the impression that such events are totally preventable leads to unrealistic expectations and unachievable goals."

So why am I bringing this up again?

Take a look at a few of the comments from the RT.com story [printed verbatim]:

Old news, as many as a million die each year cause of doctor errors. Thats why their malpractice insurance is so high. Legal unintentional homicide.

It's convenient to claim such deaths are errors but a great many are deliberate. They know such incidents will not be investigated as crimes. It's very easy to conceal a murder if no one is looking. The medical system is completely corrupt.

if they'd stop getting high in med school and pay more attention maybe this wouldnt happen. then there is their attitudes. Heaven forbid anyone needs medical care, that's for sure.

According to CDC, medical errors is not even a category of death, but they published research that indicates drunk drivers kill about 10,000 yearly. If that is correct, then doctors kill almost twice that many every hour of every day -. MADD should be mad about DEADLY DOCTORS. You are 40 times more likely to be killed by a deadly doc than you are by a drunk driver. And yet - where is the "funding" for this deadly phenomena?

I know those who comment on the Internet usually do not represent the views of rational individuals, but it infuriates the hell out of me that the 440,000 deaths from medical errors estimate, which is clearly wrong, is repeatedly trumpeted all over the place and so readily believed.

By the way, the paper appeared in the Journal of Patient Safety, which recently underwent an editorial change due to a kickback scandal involving former editor Dr. Charles Denham. That's another story (here).

Do doctors and hospitals make mistakes? Yes. Can we improve? Yes. Does it help to exaggerate the magnitude of the problem? Emphatically, no.

Thursday, February 20, 2014

Single-incision vs. standard 4-port laparoscopic cholecystectomy: Part 2

Here's another paper that shows why reading only an abstract can sometimes be misleading.

A prospective trial (abstract here) of 49 patients randomized to single-incision laparoscopic cholecystectomy (SILC) vs. 51 who had standard 4-port laparoscopic cholecystectomy (LC) found that average operative times were 63.5 ± 21.0 minutes for the SILC compared to 43.8 ± 24.2 minutes for those who had LC, and hospital charges were also more than $4000 higher for the SILC patients—both significant differences with p values < 0.0001.

Medical and surgical supplies were the major factors contributing to the increased charges for SILC.

Other than a significantly larger number of females in the SILC group, the patients were similar in baseline characteristics.

Other important considerations such as postoperative pain, hospital length of stay (an average of 24 hours or less for both operations), use of analgesics, cosmetic appearance of the wounds, rates of incisional hernia, and quality of life were similar. Average follow-up was 16 months in both groups. The authors concluded that there was no advantage to SILC.

Since this paper supports my bias against single-incision surgery, I was going to tout it as yet another negative paper like a recent meta-analysis (here) from a group in Croatia showing absolutely no advantage for SILC.

But this sentence from the "Methods" section of the paper foiled my plan. "Before partaking in the study, each surgeon developed his or her SILC technical skills in a laboratory setting and demonstrated proficiency during 5 SILCs under the supervision of a surgeon with experience on more than 50 SILC cases."

This was not mentioned in the abstract.

Are you surprised that a surgeon that might take longer to do SILC, an operation done only 5 times before, than LC, which each of the surgeons had probably done hundreds of times? Although the mean operative duration was longer for SILC, it is a "straw man" in statistical parlance. This may not detract from the rest of the results but certainly has to be considered.

As noted by the authors, the study was underpowered (that is, there weren't enough patients) to detect differences in some of the other outcomes due to difficulty recruiting subjects.

Of 946 patients offered enrollment in the study, only 103 consented. Patients declined to participate either because the surgeon explained that he had done more standard LC procedures, or the patients opted for the SILC because of its supposed cosmetic advantage.

The authors, based at Northwestern University Medical School in Chicago, should be commended for their honesty in explaining their inexperience with SILC to potential subjects of the trial and wonder if other surgeons who perform SILC do this.

This paper also highlights the problems associated with attempts to conduct randomized prospective studies involving new surgical procedures.

Bottom line: The extra costs associated with SILC are not worth it.

Part 1 of this 2-part series on SILC appeared on Tuesday, 2/18.

Tuesday, February 18, 2014

Single-incision vs. standard 4-port laparoscopic cholecystectomy: Part 1 of 2



The saying used to be, "You can get any paper published if you have enough stamps." Now with electronic submission, you don't even need the stamps.

A retrospective study comparing single-incision laparoscopic cholecystectomy (SILC) to standard 4-port laparoscopic cholecystectomy (LC) concluded that "SILC showed no disadvantage concerning risk profiles, operative times or hospital stay."

According to the abstract, 81.7% of the 115 SILC patients had elective surgery vs. 55.5% of the 344 in the LC group. The SILC cohort experienced significantly shorter operative times (70 ± 31 vs. LC: 80 ± 27 minutes) and hospital lengths of stay (3.02 ± 1.4 vs. LC: 4.6 ± 2.8 days), p < 0.001 for both. LC was converted to open surgery in 21 cases vs. none of the SILCs, p= 0.003. Rates of bile leak and incisional hernia did not differ.

Do you see any problems with this study? I do.

The groups were not really comparable because the LC group underwent more emergency operations. That difference is significant with a p value of 0.007—conveniently omitted from the abstract. The preponderance of elective cases likely accounts for the SILC group's shorter operative duration, lower rate of conversion to open, and shorter length of stay. The SILC patients were also a mean of 10 years younger.

The average operative time for the LC patients, 80 minutes, is much longer than the 40 to 45 minutes reported in most other recent series such as this one. In statistical circles, measuring one's pet theory against a false comparator is known as setting up a "straw man." I've written about this before.

This study was done in Germany, where the hospital lengths of stay for both types of surgery are far longer than those seen in the United States where about 90% of patients go home within 24 hours of laparoscopic cholecystectomy.

The authors concluded that "SILC can be regarded as a natural evolution in the era of minimally invasive surgery."

On the other hand "No disadvantage" is another way of saying, "No advantage."

This paper didn't convince me about the value of SILC. How about you?

Part 2 of this 2-part series on SILC appeared on Thursday, 2/20.