[Gluster-users] How to diagnose volume rebalance failure?

PuYun cloudor at 126.com
Mon Dec 14 13:51:43 UTC 2015


Hi,

Thank you for your reply. I don't know how to send you the huge sized rebalance log file which is about 2GB. 

However, I might have found out the reason why the task failed. My gluster server has only 2 cpu cores and carries 2 ssd bricks. When the rebalance task began, top 3  processes are 70%~80%, 30%~40 and 30%~40 cpu usage. Others are less than 1%. But after a while, 2 CPU cores are used up totally and I even can't login until the rebalance task failed. 

It seems 2 bricks require 4 CPU cores at least. Now I upgrade the virtual server with 8 CPU cores and start rebalance task again. Everything goes well for now.

I will report again when the current task completed or failed.



PuYun
 
From: Nithya Balachandran
Date: 2015-12-14 18:57
To: PuYun
CC: gluster-users
Subject: Re: [Gluster-users] How to diagnose volume rebalance failure?
Hi,
 
Can you send us the rebalance log?
 
Regards,
Nithya
 
----- Original Message -----
> From: "PuYun" <cloudor at 126.com>
> To: "gluster-users" <gluster-users at gluster.org>
> Sent: Monday, December 14, 2015 11:33:40 AM
> Subject: Re: [Gluster-users] How to diagnose volume rebalance failure?
> 
> Here is the tail of the failed rebalance log, any clue?
> 
> [2015-12-13 21:30:31.527493] I [dht-rebalance.c:2340:gf_defrag_process_dir]
> 0-FastVol-dht: Migration operation on dir
> /for_ybest_fsdir/user/Weixin.oClDcjhe/Ny/5F/1MsH5--BcoGRAJPI took 20.95 secs
> [2015-12-13 21:30:31.528704] I [dht-rebalance.c:1010:dht_migrate_file]
> 0-FastVol-dht:
> /for_ybest_fsdir/user/Weixin.oClDcjhe/Kn/hM/oHcPMp4hKq5Tq2ZQ/flag_finished:
> attempting to move from FastVol-client-0 to FastVol-client-1
> [2015-12-13 21:30:31.543901] I [dht-rebalance.c:1010:dht_migrate_file]
> 0-FastVol-dht:
> /for_ybest_fsdir/user/Weixin.oClDcjhe/PU/ps/qUa-n38i8QBgeMdI/userPoint:
> attempting to move from FastVol-client-0 to FastVol-client-1
> [2015-12-13 21:31:37.210496] I [MSGID: 109081]
> [dht-common.c:3780:dht_setxattr] 0-FastVol-dht: fixing the layout of
> /for_ybest_fsdir/user/Weixin.oClDcjhe/Ny/7Q
> [2015-12-13 21:31:37.722825] I [MSGID: 109045]
> [dht-selfheal.c:1508:dht_fix_layout_of_directory] 0-FastVol-dht: subvolume 0
> (FastVol-client-0): 1032124 chunks
> [2015-12-13 21:31:37.722837] I [MSGID: 109045]
> [dht-selfheal.c:1508:dht_fix_layout_of_directory] 0-FastVol-dht: subvolume 1
> (FastVol-client-1): 1032124 chunks
> [2015-12-13 21:33:03.955539] I [MSGID: 109064]
> [dht-layout.c:808:dht_layout_dir_mismatch] 0-FastVol-dht: subvol:
> FastVol-client-0; inode layout - 0 - 2146817919 - 1; disk layout -
> 2146817920 - 4294967295 - 1
> [2015-12-13 21:33:04.069859] I [MSGID: 109018]
> [dht-common.c:806:dht_revalidate_cbk] 0-FastVol-dht: Mismatching layouts for
> /for_ybest_fsdir/user/Weixin.oClDcjhe/Ny/7Q, gfid =
> f38c4ed2-a26a-4d83-adfd-6b0331831738
> [2015-12-13 21:33:04.118800] I [MSGID: 109064]
> [dht-layout.c:808:dht_layout_dir_mismatch] 0-FastVol-dht: subvol:
> FastVol-client-1; inode layout - 2146817920 - 4294967295 - 1; disk layout -
> 0 - 2146817919 - 1
> [2015-12-13 21:33:19.979507] I [MSGID: 109022]
> [dht-rebalance.c:1290:dht_migrate_file] 0-FastVol-dht: completed migration
> of
> /for_ybest_fsdir/user/Weixin.oClDcjhe/Kn/hM/oHcPMp4hKq5Tq2ZQ/flag_finished
> from subvolume FastVol-client-0 to FastVol-client-1
> [2015-12-13 21:33:19.979459] I [MSGID: 109022]
> [dht-rebalance.c:1290:dht_migrate_file] 0-FastVol-dht: completed migration
> of /for_ybest_fsdir/user/Weixin.oClDcjhe/PU/ps/qUa-n38i8QBgeMdI/userPoint
> from subvolume FastVol-client-0 to FastVol-client-1
> [2015-12-13 21:33:25.543941] I [dht-rebalance.c:1010:dht_migrate_file]
> 0-FastVol-dht:
> /for_ybest_fsdir/user/Weixin.oClDcjhe/PU/ps/qUa-n38i8QBgeMdI/portrait_origin.jpg:
> attempting to move from FastVol-client-0 to FastVol-client-1
> [2015-12-13 21:33:25.962547] I [dht-rebalance.c:1010:dht_migrate_file]
> 0-FastVol-dht:
> /for_ybest_fsdir/user/Weixin.oClDcjhe/PU/ps/qUa-n38i8QBgeMdI/portrait_small.jpg:
> attempting to move from FastVol-client-0 to FastVol-client-1
> 
> 
> Cloudor
> 
> 
> 
> From: Sakshi Bansal
> Date: 2015-12-12 13:02
> To: 蒲云
> CC: gluster-users
> Subject: Re: [Gluster-users] How to diagnose volume rebalance failure?
> In the rebalance log file you can check the file/directory for which the
> rebalance has failed. It can mention what was the fop for whihc the failure
> happened.
> 
> _______________________________________________
> Gluster-users mailing list
> Gluster-users at gluster.org
> http://www.gluster.org/mailman/listinfo/gluster-users
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://www.gluster.org/pipermail/gluster-users/attachments/20151214/53ca3629/attachment.html>


More information about the Gluster-users mailing list